Data query method, query engine, device, medium, and program product
By reading global statistical information in the data lake to determine the query strategy, the problem of slow query speed in the data lake is solved, and faster query response and efficiency improvement are achieved.
Patent Information
- Application Number
- PCT/IB2025/052312
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-04
- Publication Date
- 2025-09-25
AI Technical Summary
How to improve the query speed of data in the data lake, especially in large-scale and diversified data storage architecture, existing technologies are difficult to effectively improve query efficiency.
By reading the global statistical information of the data lake, determining the query strategy based on the global statistical information, and directly responding to query requests, the amount of data read is reduced to improve query speed.
By using global statistical information to determine the query strategy, the amount of data read by the query device is reduced, the query strategy determination speed and the query request response speed are improved, and the delay in the local statistical information merging process is avoided.
Smart Images

Figure IB2025052312_25092025_PF_FP_ABST
Abstract
Description
[0001] Data query method, query engine, device, medium and program product technical field
[0002]
[0001] The present disclosure relates to the field of data lake technology, and more particularly to a data query method, query engine, device, medium, and program product.
[0003]
[0002] A data lake is an architecture for storing and processing large-scale, diverse data. Similar to a database, a data lake can also store data in the form of tables. In practice, a query engine can query the data on the lake using statistical information on the data on the lake. The statistical information can be collected after the data is stored on the lake.
[0004]
[0003] Based on the above description, how to improve the query speed of lake data becomes a problem to be solved urgently.
[0005]
[0004] In view of this, the embodiments of the present disclosure provide a data query method, query engine, device, medium and program product to improve the query speed of lake data.
[0006]
[0005] In a first aspect, an embodiment of the present disclosure provides a data query method, comprising: in response to a query request of a data lake, reading global statistical information of the data lake; determining a query strategy corresponding to the query request based on the global statistical information; and responding to the query request according to the query strategy.
[0007]
[0006] In a second aspect, an embodiment of the present disclosure provides a query engine, comprising: a processing component and a query optimization component; the processing component is configured to read global statistical information of the data lake in response to a query request of the data lake; respond to the query request according to a query strategy determined by the query optimization component; and the query optimization component is configured to determine a query strategy corresponding to the query request based on the global statistical information.
[0008]
[0007] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a memory for storing one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the data query method described in the first aspect. The electronic device may also include a communication interface for communicating with other devices or communication systems.
[0009]
[0008] In a fourth aspect, embodiments of the present disclosure provide a non-transitory machine-readable storage medium having executable code stored thereon. When the executable code is executed by a computing system of an electronic device, the computing system is enabled to implement at least the data query method described in the first aspect.
[0009] In a fifth aspect, embodiments of the present disclosure provide a computer program product. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is enabled to implement the data query method described in the first aspect.
[0010]
[0010] The data query method provided in the embodiments of the present disclosure can first read the global statistical information of the data lake after initiating a query request to the data lake. Then, the query strategy corresponding to the query request is determined based on the global statistical information, and finally the query request is responded to according to the query strategy. The global statistical information is the statistical result obtained by analyzing all data in the data lake and is used to describe the overall characteristics of the data in the data lake. Correspondingly, the local statistical information of the data lake is the statistical information obtained by analyzing the data contained in the data partitions in the data lake and is used to describe the local characteristics of the data in the data lake. The statistical information corresponding to different data partitions can be collectively referred to as the local statistical information of the data lake.
[0011]
[0011] As can be seen, compared to the local statistical information of the data lake, the global statistical information of the data lake contains less data. Therefore, by directly reading the global statistical information to determine the query strategy, the amount of data read can be reduced, thereby increasing the speed of determining the query strategy and also increasing the speed of querying data in the data lake.
[0012]
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0013]
[0013] FIG1 is a flow chart of a data query method provided in an embodiment of the present disclosure;
[0014] FIG2 is a flow chart of a statistical information merging method provided in an embodiment of the present disclosure;
[0015] FIG3 is a schematic diagram of a process for writing characteristic values of a data partition provided by an embodiment of the present disclosure;
[0016]
[0016] FIG4 is a schematic diagram of a data structure storing characteristic values of a data lake provided by an embodiment of the present disclosure;
[0017] FIG5 is a structural diagram of a query engine provided in an embodiment of the present disclosure;
[0018]
[0018] FIG6 is a schematic diagram of the structure of a data query device provided in an embodiment of the present disclosure;
[0019] FIG. 7 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure.
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all of them. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present disclosure without creative work are within the scope of protection of the present disclosure.
[0021]
[0021] The terms used in the embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a," "the," and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms. Unless the context clearly indicates otherwise, "a plurality" generally includes at least two, but does not exclude the inclusion of at least one.
[0022]
[0022] It should be understood that the term "and / or" used herein is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and A exists alone.
[0023]
[0023] As used herein, the words “if” or “when” may be interpreted as “upon” or “upon” or “in response to determining” or “in response to identifying” depending on the context. Similarly, the phrases “if it is determined” or “if (a stated condition or event) is identified” may be interpreted as “upon determination” or “in response to determining” or “upon identifying (a stated condition or event)” or “in response to identifying (a stated condition or event)” depending on the context.
[0024]
[0024] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0025]
[0025] It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or system. In the absence of further limitations, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the product or system comprising the element.
[0026]
[0026] Before describing the methods and engines provided by the following embodiments of the present disclosure, the following concepts may also be explained.
[0027] Data lake: An architecture for storing and processing large-scale, diverse data. Within a data lake, data can be stored in tables. Compared to databases, data lakes have lower requirements for data structure.
[0027]
[0028] Data partition: A table in a data lake can be divided into at least one sub-table. Each sub-table is a data partition. A data partition in a data lake contains at least one field.
[0028]
[0029] Local statistics of a data lake: Statistics are collected on a field-by-field basis across the data in any data partition within the data lake. These statistics are considered the statistics for that data partition. A data lake can contain multiple data partitions; the statistics for each partition are collectively referred to as the local statistics for the data lake. Local statistics can be used to describe the local characteristics of the data within the data lake. Global statistics of a data lake: Statistics are collected on a field-by-field basis across all data contained in all data partitions within the data lake. These statistics are considered the global statistics for the data lake. Global statistics can be used to describe the overall characteristics of the data within the data lake.
[0029]
[0030] Optionally, both local and global statistical information may include at least one of the following: the total number of rows, the amount of data, the maximum value in a field, the null ratio (Null Ratio) of a field, and the number of distinct values (NDV) in a field, that is, the field cardinality of the field.
[0030]
[0031] Global statistics can include: the total number of rows, data volume, maximum value, null value ratio, field cardinality, etc. for each field in all data partitions within the data lake. Local statistics can also include: the total number of rows, data volume, maximum value, null value ratio, field cardinality, etc. for each field in each data partition within the data lake.
[0031]
[0032] As can be seen, global statistical information can be a set of statistical data obtained by analyzing all data in the data lake. This set of statistical data can include at least one of the items mentioned above. Local statistical information can be a set of statistical data for any data partition in the data lake, obtained by analyzing the data contained in that data partition. This set of statistical data can also include at least one of the items mentioned above.
[0032]
[0033] In practice, since a data lake contains multiple data partitions, the local statistical information of the data lake can include a set of statistical information for each of the multiple data partitions, that is, it includes multiple sets of statistical data. Therefore, the amount of data in the global statistical information is much smaller than that in the local statistical information.
[0033]
[0034] Query Engine: A component used to query data in the data lake. The query engine can also perform statistics on the data in the data lake to obtain local and global statistical information for the data lake.
[0034]
[0035] Based on the above introduction, some embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. The following embodiments and features may be combined unless they conflict with each other. Furthermore, the sequence of steps in the following method embodiments is provided for illustrative purposes only and is not intended to be a strict limitation.
[0035]
[0036] Figure 1 is a flowchart of a data query method provided by an embodiment of the present disclosure. This method provided by an embodiment of the present disclosure can be executed by a query device. More specifically, the method can be executed by a query engine deployed in the query device. As shown in Figure 1, the method can include the following steps.
[0036]
[0037] S101, in response to a query request of a data lake, read global statistical information of the data lake.
[0037]
[0038] S102: Determine a query strategy corresponding to the query request according to the global statistical information.
[0038]
[0039] S103: Respond to the query request according to the query strategy.
[0039]
[0040] A user can generate a query request for the data lake. In response to this request, the query device can first read the global statistical information of the data lake. Then, based on this global statistical information, it can determine the query strategy corresponding to the query request. Finally, the query device can execute this query strategy to respond to the query request.
[0040]
[0041] As described above, global statistics can be the statistical results obtained by analyzing all data in a data lake. Correspondingly, local statistics of a data lake can be the statistical results obtained by analyzing the data contained in a data partition within the data lake. The statistics corresponding to different data partitions in a data lake can be collectively referred to as the local statistics of the data lake. In practice, due to the large number of data partitions in a data lake, the data volume of the local statistics of a data lake is much larger than the data volume of the global statistics of the data lake.
[0041]
[0042] As described above, both local and global statistical information can include at least one of the following: the total number of rows in a column (i.e., a field), the amount of data, the maximum value, the null ratio, and the number of distinct values (NDV) in a column, also known as the field cardinality.
[0042]
[0043] Optionally, the query device can merge the statistical information of different data partitions in the data lake and determine the merged result as the global statistical information of the data lake. Optionally, for any data partition in the data lake, the statistical information of the data partition, that is, the local statistical information of the data lake, can be generated by the query device.
[0043]
[0044] Alternatively, the query device can use different operators to determine a query strategy that includes a table join order and data propagation method. The table join order can be determined by the Join operator, and the data propagation method can be determined by the Exchange operator. The execution costs of the operators can then be estimated based on global statistical information from the data lake. Ultimately, the table join order and data propagation method with the lowest execution cost can be determined, thereby determining the query strategy.
[0044]
[0045] Next, the query device joins the multiple tables involved in the query request according to the table join order specified in the query strategy to obtain candidate data. It then uses the query conditions specified in the query request to obtain the target data that meets the query conditions from the candidate data. The target data is then fed back to the user as a response to the query request using the data propagation method specified in the query strategy, thus completing the query request response.
[0045]
[46] The generation and merging methods of local statistical information can be referred to the description in the following related embodiments. Optionally, the merging of statistical information can be performed periodically.
[0046]
[47] The following example illustrates the relationship between global and local statistics.
[0047]
[48] Assume that the data lake contains three data partitions, namely data partition 1 to data partition 3, and the data lake can store data of the two fields "user name" and "height": [User 1, height 150cm], [User 2, height
[0048] 170cm], [User 3, height 160cm], [User 4, height 150cm], [User 5, height 180cm], [User 6, height 175cm], [User 7, height 180cm], [User 8, height 185cm], [User
[0049] 9, height 160cm].
[0050]
[49] The above data can be stored in three data partitions respectively. That is, the data contained in data partition 1 can be: [User 1, height 150cm], [User 2, height 170cm], [User 3, height 160cm]. The data contained in data partition 2 can be: [User 4, height 150cm], [User 5, height 180cm], [User 6, height 150cm]. The data contained in data partition 3 can be: [User 7, height 180cm], [User 8, height 185cm], [User 9, height 160cm].
[0051]
[50] Taking the "height" field as an example, the statistical information of data partition 1 may include: the maximum value is 170cm, the minimum value is 150cm, the total number of rows is 3, and the field cardinality is 3. The statistical information of data partition 2 may include: the maximum value is 180cm, the minimum value is 150cm, the total number of rows is 3, and the field cardinality is 2. The statistical information of data partition 3 may include: the maximum value is 185cm, the minimum value is 160cm, the total number of rows is 3, and the field cardinality is 3.
[0052]
[51] The query device can merge the statistical information of the above-mentioned "height" field, and the final global statistical information of the data lake is: the maximum value is 185cm, the minimum value is 150cm, the total number of rows is 9, and the field cardinality is 5.
[0053]
[52] In this embodiment, after initiating a query request to the data lake, the global statistical information of the data lake can be read first. Then, the query strategy corresponding to the query request is determined based on the global statistical information, and the query request is finally responded to according to the query strategy. Since the global statistical information of the data lake contains less data than the local statistical information of the data lake, determining the query strategy by directly reading the global statistical information can reduce the amount of data read by the query device, thereby increasing the speed of determining the query strategy and further improving the efficiency of querying data in the data lake.
[0054]
[53] In addition, the technical effects that can be achieved by the data query method provided in each embodiment of the present disclosure can also be understood in conjunction with the following content.
[0055]
[54] In practice, local statistical information of the data lake can also be used to respond to query requests. The specific process can be as follows:
[0055] In response to a query request, the query device can first obtain target statistical information related to the query request from the local statistical information of the data lake, and then merge the target statistical information. Then, the query device can determine a query strategy based on the merged result and respond to the query request according to the query strategy. The target statistical information can be statistical information of the target data partition related to the query request.
[0056] When using the above approach, on the one hand, after a query request is generated, the query device needs to determine and merge the target statistical information in real time upon receiving the query request. This real-time statistical process obviously slows down the query strategy determination speed, which can further slow down the query request response speed. On the other hand, the more target data partitions associated with the query request, the longer it takes for the query device to merge the statistical information, which also affects the query request response speed.
[0057]
[0057] Compared to the above approach, the disclosed embodiments propose global statistical information that can be used in data lake scenarios. Furthermore, with the help of this global statistical information, the query device can omit the process of merging local statistical information during the query request response process, thereby improving the speed of determining the query strategy and further improving the query request response speed. Furthermore, because the query device directly determines the query strategy based on the global statistical information, the query request response speed is not affected regardless of the number of target data partitions involved in the query request. In other words, the use of global statistical information ensures that the number of target data partitions does not affect the query request response speed.
[0058]
[0058] Optionally, the statistical information of the data lake can be stored in a database (meta store) as metadata of the data in the data lake. Optionally, the database can be independent of the data lake.
[0059]
[0059] Optionally, when the database stores global statistical information of the data lake, the query device can also utilize its own cache mechanism to improve the response speed of the query request. Specifically, after receiving a query request, the query device can preferentially read the global statistical information of the data lake from the local cache. If the query device's local cache does not contain the global statistical information, the query device reads the global statistical information from the database and stores it in the local cache.
[0060]
[0060] Based on the cache mechanism, the technical effect that can be achieved by using global statistical information to perform data query can also be understood in conjunction with the following content.
[0061]
[0061] Since the global statistical information has a smaller data volume and occupies a smaller storage space than the local statistical information, the global statistical information can be stored more completely in the cache of the query device than the local statistical information. This can improve the hit rate of the information in the cache during the query request response process. That is, when responding to the query request, the query device is likely to read the global statistical information from the cache, thereby also improving the response speed of the query request.
[0062]
[0062] Optionally, the database can also be shared by multiple data lakes. In this case, the technical effects that can be achieved by using global statistical information for data query can also be understood in conjunction with the following content.
[0063] When a database is shared by multiple data lakes and stores local statistical features of the data lakes, and a query request is directed to a target data lake among the multiple data lakes and the number of target data partitions associated with the query request is large, the query device corresponding to the target data lake needs to read a large amount of local statistical information from the database when responding to the query request. Reading a large amount of local statistical information consumes a significant amount of resources and may prevent query devices corresponding to other data lakes from properly reading the local statistical information of other data lakes from the database, thereby rendering the other data lakes inoperable.
[0064]
[0064] In the embodiments provided by the present disclosure, even if the number of target data partitions related to the query request is large, the query device can directly use the global statistical information of the target data lake with a small amount of data read from the database to respond to the query request. Therefore, the above-mentioned situation of affecting the normal use of other data lakes will not occur.
[0065]
[0065] In the embodiment shown in FIG. 1 , it has been mentioned that the query device can obtain global statistical information of the data lake by merging local statistical information. Furthermore, based on the example in the embodiment shown in FIG. 1 , it can be seen that the total number of rows and the maximum value in the statistical information can be directly merged. That is, the total number of rows in the local statistical features can be directly added to obtain the total number of rows in the global statistical information. The maximum value in the local statistical information can be compared to obtain the maximum value in the global statistical information.
[0066]
[0066] Obviously, the field cardinality in the statistical information cannot be directly obtained through simple calculation. In this case, the following embodiment can be used to merge the field cardinality in the local statistical features. FIG2 is a flow chart of a statistical information merging method provided by an embodiment of the present disclosure. The method provided by an embodiment of the present disclosure can also be executed by a query device. As shown in FIG2, the method may include the following steps.
[0067]
[0067] S201: Count the characteristic values of different data partitions stored in the first data structure to obtain characteristic values of the data lake stored in the form of a second data structure.
[0068]
[0068] S202, converting the storage format of the characteristic values of the data lake from the second data structure to the first data structure.
[0069]
[0069] S203: Determine the field cardinality of the data lake according to the characteristic values of the data lake stored in the first data structure.
[0070]
[0070] Since data in the data lake is stored in partitions, data on the lake is written to a specific data partition in the data lake. After the data is written to the data partition, the query device can first obtain the characteristic values of each data partition. Furthermore, by statistically analyzing these characteristic values, the characteristic values of the data lake can be obtained. Finally, the field cardinality of the data lake is calculated based on the characteristic values of the data lake.
[0071]
[0071] The characteristic values of the data partitions may be generated by the query device when a data partition is added or deleted. The characteristic values of each data partition and the data lake may be stored in corresponding data structures, such as tables, queues, arrays, etc. The characteristic values may be considered as basic data for calculating field cardinality.
[0072]
[0072] Optionally, the characteristic values of different data partitions in the data lake can be stored in the first data structure, and the characteristic values of the data lake can be stored in the second data structure. In order to calculate the field data of the data lake, the characteristic values of the data lake can also be converted from the second data structure and stored in the first data structure.
[0073]
[0073] Optionally, the first data structure may be a one-dimensional array, the second data structure may be a two-dimensional array, and the length of the one-dimensional array is the same as the number of rows of the two-dimensional array. Optionally, the one-dimensional array may be a one-dimensional array of the int type, and the two-dimensional array may be a two-dimensional array of the int type.
[0074]
[0074] The following will first describe the process of how the query device generates a one-dimensional array storing characteristic values of any data partition.
[0075]
[0075] For any data partition, the query device can first perform a hash calculation on the data in the partition to obtain a hash calculation result. After the hash calculation, the data in the data partition is converted into binary data. The hash calculation result, which is represented as binary data, can then be truncated into a first part and a second part that serves as an index. The maximum number of leading zeros in the first part of the hash calculation result can be written to the position pointed to by the second part of the hash calculation result in the one-dimensional array to obtain the one-dimensional array corresponding to the data partition. Wherein, the hash calculation result can be P-bit binary data, then the first part can be L bits and the second part can be T bits, where P = L + T. A common setting is that the binary data is 64 bits, L = 54, and T = 10.
[0076] The generation process of the above-mentioned one-dimensional array can also be understood in conjunction with FIG3. As shown in FIG3, assuming that the length of the one-dimensional array corresponding to the data partition is M, the elements in the one-dimensional array can be respectively referred to as delta[0] to delta[M1]. Moreover, the one-dimensional data shown in FIG3 is actually a HyperLogLog data structure.
[0077]
[0077] The query device may first perform a hash calculation on a data in the data partition. Assume that the hash calculation result of the data is 001010010110. The hash calculation result may be truncated into a first part "00101101" and a second part "00101" as an index. The number of leading zeros in the first part is 2, and the second part is converted to a decimal value of 5. "5" may be used as an index to point to delta[5] in the one-dimensional array, and the number of leading zeros 2 may be written to delta[5] in the one-dimensional array.
[0078]
[0078] The hash calculation result of another data in any data partition can also be processed in the above manner: if the index value of the other data is also 5, and the number of leading zeros is 4 which is greater than 2, then the value on delta[5] is updated to 4. After each data in any data partition is processed as above, a one-dimensional array corresponding to the data partition can be obtained.
[0079]
[0079] Optionally, before truncating the hash calculation result, the query device may bucket the hash calculation result, that is, bucket the data in the data partition. Optionally, the bucketing process may be as follows: if the hash calculation result is within the hash value range corresponding to the bucket, the hash calculation result may be placed into the bucket. The number of buckets in any data partition is the same as the length of the one-dimensional array corresponding to the data partition. Optionally, the number of buckets may be preset, and L may be calculated based on a functional relationship between the preset number of buckets and the aforementioned L.
[0080]
[80] In practice, data belonging to the same bucket often have the same index. Therefore, the one-dimensional array corresponding to any data partition actually records the maximum number of leading zeros corresponding to the hash calculation result of the data in each bucket. This maximum number is also the characteristic value of the data partition.
[0081]
[81] In this embodiment, the characteristic value of any data partition in the data lake can be obtained by means of bucketing, hash calculation, leading zero calculation, etc. The characteristic value can be stored in an HLL data structure, that is, a one-dimensional array.
[0082]
[82] After obtaining a one-dimensional array storing the characteristic values of the data partitions based on the above method, this one-dimensional array can optionally be used to generate a two-dimensional array recording the characteristic values of the data lake. As shown in Figure 4, assuming that the characteristic values of the data lake can be stored in an M*N two-dimensional array, the meaning of the value stored at any position [i, j] in the two-dimensional array can be defined as: In the one-dimensional array corresponding to each of the different data partitions contained in the data lake, the value stored at delta[i] is the number of j.
[0083]
[83] Based on the above definition, the query device can generate a two-dimensional array corresponding to the data lake based on the one-dimensional arrays corresponding to different data partitions in the data lake. The target elements in the second array are the number of target values in the one-dimensional array. The target elements correspond to the target number of rows and the target number of columns in the second array. The position of the target value in the one-dimensional array is the same as the target number of rows, and the target value is the same as the target number of columns.
[0084]
[84] In this embodiment, the query device can provide a new data structure for storing characteristic values of the data lake. Based on this data structure, the field cardinality of the data lake can be calculated more accurately.
[0085]
[85] In addition, when data partitions in the data lake are added or deleted, the field cardinality of the data lake also needs to be updated. An optional update method can be: updating the one-dimensional array that records the characteristic values corresponding to each data partition, and then merging the updated one-dimensional arrays. Finally, the field cardinality of the data lake can be determined based on the merged result. Assuming that the data lake includes 1 million data partitions and the size of a one-dimensional array is 1K, when updating according to the above method, the amount of data that the query device needs to read is 1 million * 1K = 1G.
[0086]
[86] With the help of the two-dimensional array provided in Figure 4, another optional update method is: in response to the addition or deletion of data partitions, the query device can directly merge the one-dimensional array corresponding to the added or deleted data partition into the two-dimensional array corresponding to the data lake, that is, update the two-dimensional array corresponding to the data lake with the one-dimensional array corresponding to the added or deleted data partition. The query device can ultimately determine the field cardinality of the data lake based on the updated two-dimensional array. More specifically, the query device can, based on preset rules, add the feature values in the one-dimensional array corresponding to the added data partition to the two-dimensional array corresponding to the data lake; or delete the feature values in the one-dimensional array corresponding to the deleted data partition from the two-dimensional array corresponding to the data lake.
[0087]
[0087] Assume that the data lake includes 1 million data partitions, each data partition has 2048 buckets, the first part of the truncated binary data is 54 bits, and int type data occupies 4 bytes. Using this update method, the amount of data read by the query device during the merge process is 54 * 2048 * 4 (the size of the two-dimensional array of the data lake) + 1K (the size of the one-dimensional array of the newly added or deleted data partition) = 433K.
[0088]
[0088] This embodiment provides an incremental feature value merging method. Furthermore, in this embodiment, by updating the feature values using a two-dimensional array that records the feature values of the data lake, the amount of data read by the query device can be significantly reduced. When data partitions in the data lake are added or deleted, the updated two-dimensional array corresponding to the data lake can be obtained more quickly.
[0089]
[0089] The following will further describe in detail the updating process of the two-dimensional array when data partitions are added or deleted.
[0090]
[0090] When a new data partition is added to the data lake, in response to the addition of the data partition, the query device may obtain a one-dimensional array corresponding to the new data partition and a two-dimensional array corresponding to the data lake. The query device may then sequentially traverse the one-dimensional array corresponding to the new data partition and increment the elements corresponding to the first reference row number and the first reference column number in the two-dimensional array corresponding to the data lake by one, where the first reference column number is the same as the first element in the one-dimensional array corresponding to the newly added data partition currently traversed, and the first reference row number and the first element have the same position in the one-dimensional array corresponding to the newly added data partition.
[0091]
[0091] The above process can also be understood in conjunction with the following pseudo code: for(K->0:M-1) {gsm[K][Al.get(K)]++;
[0092] }
[0093]
[0092] Wherein, gsm is the name of the two-dimensional data, A1 is the one-dimensional array corresponding to the newly added data partition, and A1.get(K) is the data stored at the delta[K] position in the one-dimensional array.
[0094]
[0093] For example, for a one-dimensional array of length M corresponding to a newly added data partition, if K=1, that is, the characteristic value at the delta[1] position currently traversed in the one-dimensional array is W, the query device may add 1 to the value at the [1, W] position in the two-dimensional array. Then, the characteristic value at the delta[2] position in the one-dimensional array is continued to be traversed until the characteristic value at the delta[M1] position is traversed. At this point, the characteristic value of the newly added data partition is updated to the characteristic value of the data lake.
[0094] When a data partition is deleted from the data lake, in response to the deletion of the data partition, the one-dimensional array corresponding to the deleted data partition and the two-dimensional array corresponding to the data lake are obtained. Then, the query device may sequentially traverse the one-dimensional array corresponding to the deleted data partition and subtract one from the elements corresponding to the second reference row number and the second reference column number in the two-dimensional array corresponding to the data lake, wherein the second reference column number is the same as the second element in the one-dimensional array corresponding to the currently traversed deleted data partition, and the second reference row number is the same as the position of the second element in the one-dimensional array corresponding to the deleted data partition.
[0095]
[0095] The above process can also be understood in conjunction with the following pseudo code: for(K->0:M1) {gsm[K][A2.get(K)] -;
[0096] )
[0097]
[0096] Wherein, gsm is the name of the two-dimensional data, A2 is the one-dimensional array corresponding to the deleted data partition, and A2.get(K) is the data stored at the delta[K] position in the one-dimensional array.
[0098]
[0097] Continuing with the example, for a one-dimensional array of length M corresponding to the deleted data partition, if K=1, that is, the feature value at the delta[1] position in the one-dimensional array currently traversed is W, the query device can add 1 to the value at the [1, W] position in the two-dimensional array. Then, the query device continues to traverse the feature values at the delta(2) position in the one-dimensional array until it reaches the feature value at the delta[M-1] position. At this point, the feature value of the deleted data partition is completely deleted from the feature values of the data lake.
[0099]
[0098] In addition, according to the embodiment shown in FIG2 , in order to realize the calculation of the field data of the data lake, the characteristic values of the data lake also need to be converted and stored from the second data structure to the first data structure. More specifically, the storage format of the characteristic values of the data lake is converted from a two-dimensional array to a one-dimensional array.
[0100]
[0099] In an optional conversion method, the query device can sequentially traverse different rows in the two-dimensional array corresponding to the data lake in a preset direction to obtain the first non-zero data in each row. According to the order in which the non-zero data corresponding to different rows in the two-dimensional array are obtained, the non-zero data corresponding to different rows in the two-dimensional array are sequentially written into the one-dimensional array corresponding to the data lake according to the column number of the non-zero data in the two-dimensional array.
[0101]
[0100] Optionally, the preset direction may be to traverse forward from the last column of the two-dimensional array. Continuing with the above example, for a one-dimensional array of length M and a two-dimensional array of length M*N, and assuming that traversal can be performed row by row starting from row 0 of the two-dimensional array, it can be determined through traversal that the first non-zero data in row 0 is in column 5. Then, the column number "5" of the non-zero data can be written into delta[0] of the one-dimensional array.
[0102] Similarly, by traversing, it can be determined that the first non-zero data in the first row of the two-dimensional array is on the 8th column, and then the column number "8" of the non-zero data can be written into delta[1] of the one-dimensional array. Similarly, when the first non-zero data on the M1th row in the two-dimensional array is on the 6th column, the column number "6" can be written into delta[M1] of the one-dimensional array. The above conversion process can also be understood in conjunction with the following pseudo code: for(k->0:M1) { for(i->N1:0) { if(gsm[k][i]>0) {
[0103] R.set(k,i); break;
[0104]
[0102] Wherein, gsm is the name of the two-dimensional data, and ML N-1 are the maximum number of rows and the maximum number of columns of the two-dimensional array respectively.
[0105]
[0103] In this embodiment, the second array storing the characteristic values of the data lake can be converted into the first array, which can make the field cardinality calculated subsequently based on the characteristic values more accurate.
[0106]
[0104] The working process of the query device has been described above from a method perspective. The above process can be specifically performed by a query engine in the query device. The following also describes the table file merging process from the perspective of the query engine. FIG5 is a schematic diagram of the structure of a query engine provided in an embodiment of the present disclosure. As shown in FIG5, the query engine may include: a processing component and a query optimization component.
[0107]
[0105] In response to a query request from the data lake, the processing component may read global statistical information from the data lake. The query optimization component may determine a query strategy corresponding to the query request based on the global statistical information, and the query strategy may be sent to the processing component. Ultimately, the processing component may respond to the query request according to the query strategy determined by the query optimization component. Optionally, the processing component may combine the statistical information of different data partitions in the data lake, i.e., the local statistical information of the data lake, to obtain global statistical information of the data lake.
[0108]
[0106] In this embodiment, after a query request is initiated to a data lake, the processing component may first read the global statistical information of the data lake. Then, the query optimization component may determine a query strategy corresponding to the query request based on the global statistical information. Ultimately, the processing component may respond to the query request according to this query strategy. Because the global statistical information of the data lake contains less data than the local statistical information of the data lake, the query optimization component uses the global statistical information read by the processing component to determine the query strategy. This can reduce the amount of data read by the processing component, thereby increasing the speed of query strategy determination and further improving the efficiency of querying data in the data lake.
[0107] In addition, compared to using the global statistical information of the data lake to respond to query requests, in the embodiment shown in FIG. 5 , when the query engine uses the global statistical information of the data lake to respond to query requests, it can omit the process of merging the target statistical information of the target data partition corresponding to the query request, thereby increasing the speed of query strategy determination and further improving the speed of query request response. Furthermore, because the query engine determines its query strategy directly based on global statistical information, the query response speed is not affected by the number of target data partitions involved in the query request. In other words, the use of global statistical information ensures that the number of target data partitions does not affect the query response speed.
[0109]
[0108] In addition, for the contents not described in detail in this embodiment and the technical effects that can be achieved, please refer to the relevant descriptions in the above embodiments, and no further details will be given here.
[0110]
[0109] Optionally, the statistical information of the data lake can be stored in a database (meta store) as metadata of the data in the data lake. Optionally, the database can be independent of the data lake.
[0111]
[0110] Optionally, when the database stores global statistical information of the data lake, the query device can also use its own cache mechanism to improve the response speed of the query request. For details, please refer to the relevant description in the above embodiment and will not be repeated here.
[0112]
[0111] In addition, optionally, the database can also be shared by multiple data lakes. In this case, the technical effect that can be achieved by using global statistical information to perform data query can also be understood in conjunction with the following content:
[0113] When a database is shared by multiple data lakes and stores local statistical features of the data lakes, and a query request is directed to a target data lake among the multiple data lakes and the number of target data partitions associated with the query request is large, the query device can directly use the global statistical information of the target data lake, which has a small amount of data, read from the database to respond to the query request. Therefore, the aforementioned situation of affecting the normal use of other data lakes will not occur. For related details, please refer to the description of the above embodiment and will not be repeated here.
[0114]
[0113] Optionally, the embodiment shown in FIG. 5 has already mentioned that the processing component in the query engine can obtain global statistical information of the data lake by merging local statistical information. Furthermore, as shown in the example of the embodiment shown in FIG. 1 , the total number of rows and the maximum value in the statistical information can be directly merged. However, the field cardinality in the statistical information cannot be directly merged. The processing component can also obtain the field cardinality in the global feature information using the method shown in the embodiments shown in FIG. 2 to FIG. 4 .
[0115]
[0114] The feature data of the data partitions used in the field cardinality merging process may also be generated by a processing component in the query engine, and the field cardinality update may also be performed by the processing component. The specific process can be found in the description of the above-mentioned related embodiments and will not be repeated here.
[0115] In addition, for the contents not described in detail in this embodiment and the technical effects that can be achieved, please refer to the relevant description of the above-mentioned embodiments and will not be repeated here.
[0116]
[0116] The data query device of one or more embodiments of the present disclosure will be described in detail below. Those skilled in the art will appreciate that these data query devices can be configured using commercially available hardware components through the steps taught in this solution.
[0117]
[0117] FIG6 is a structural diagram of a data query device provided by an embodiment of the present disclosure. As shown in FIG6, the device may include the following modules.
[0118]
[0118] A reading module 11 is configured to read global statistical information of the data lake in response to a query request of the data lake.
[0119]
[0119] The strategy determination module 12 is used to determine the query strategy corresponding to the query request based on the global statistical information.
[0120]
[0120] A response module 13 is configured to respond to the query request according to the query strategy.
[0121]
[0121] The method is applied to a query engine, the global statistical information of the data lake is stored in a database shared by multiple data lakes, or stored in a cache of a query device, and the query engine is deployed in the query device.
[0122]
[0122] Optionally, the device further includes: a global statistical information determination module 14, configured to determine the global statistical information of the data lake based on the respective statistical information of different data partitions in the data lake.
[0123]
[0123] Optionally, the statistical information of the data partition includes the field cardinality of the data partition, and the global statistical information includes the field cardinality of the data lake.
[0124]
[0124] The global statistical information determination module 14 is configured to collect statistics on the characteristic values of the different data partitions stored in the first data structure to obtain characteristic values of the data lake stored in the form of a second data structure; convert the storage form of the characteristic values of the data lake from the second data structure to the first data structure; and determine the field cardinality of the data lake based on the characteristic values of the data lake stored in the first data structure.
[0125]
[0125] Optionally, the first data structure includes a one-dimensional array, the second data structure includes a two-dimensional array, and the length of the one-dimensional array is the same as the number of rows of the two-dimensional array.
[0126]
[0126] The global statistical information determination module 14 is configured to generate a two-dimensional array corresponding to the data lake based on the one-dimensional arrays corresponding to different data partitions in the data lake, wherein the target elements in the second array are the number of target values in the one-dimensional array, the target elements correspond to the target number of rows and the target number of columns of the second array, the position of the target value in the one-dimensional array is the same as the target number of rows, and the target value is the same as the target number of columns.
[0127] Optionally, the apparatus further includes: a local statistical information determination module 15, configured to perform a hash calculation on data in any data partition in the data lake; truncate the hash calculation result into a first part and a second part serving as an index, where the number of columns in the two-dimensional array is the same as the number of digits in the first part of the hash calculation result; and write the maximum number of leading zeros in the first part of the hash calculation result into a position pointed to by the second part of the hash calculation result in the one-dimensional array, to obtain a one-dimensional array corresponding to the any data partition.
[0127]
[0128] Optionally, the global statistical information determining module 14 is configured to sequentially traverse different rows in the two-dimensional array in a preset direction to obtain the first non-zero data in the different rows; and sequentially write the non-zero data corresponding to the different rows in the two-dimensional array by column number in the two-dimensional array into the one-dimensional array corresponding to the data lake according to the order in which the non-zero data corresponding to the different rows in the two-dimensional array are obtained.
[0128]
[0129] Optionally, the apparatus further includes: an updating module 16, configured to update a two-dimensional array corresponding to the data lake in response to additions and deletions of data partitions in the data lake.
[0129]
[0130] Optionally, the update module 16 is configured to, in response to the addition of a new data partition, obtain a one-dimensional array corresponding to the new data partition and a two-dimensional array corresponding to the data lake; traverse the one-dimensional array corresponding to the new data partition; and increment the elements corresponding to the first reference row number and the first reference column number in the two-dimensional array corresponding to the data lake by one, wherein the first reference column number is the same as the first element in the one-dimensional array corresponding to the newly added data partition currently traversed, and the first reference row number and the first element have the same position in the one-dimensional array corresponding to the newly added data partition.
[0130]
[0131] Optionally, the update module 16 is configured to, in response to deletion of a data partition, obtain a one-dimensional array corresponding to the deleted data partition and a two-dimensional array corresponding to the data lake; traverse the one-dimensional array corresponding to the deleted data partition; and subtract one from the elements corresponding to the second reference row number and the second reference column number in the two-dimensional array corresponding to the data lake, where the second reference column number is the same as the second element in the one-dimensional array corresponding to the currently traversed deleted data partition, and the second reference row number and the second element have the same position in the one-dimensional array corresponding to the deleted data partition.
[0131]
[0132] The apparatus shown in FIG6 can execute the methods of the embodiments shown in FIG1 through FIG4 . For portions not described in detail in this embodiment, reference can be made to the relevant descriptions of the embodiments shown in FIG1 through FIG4 . The implementation process and technical effects of this technical solution are described in the embodiments shown in FIG1 through FIG4 , and will not be further elaborated here.
[0132]
[0133] In one possible design, the data query methods provided in the above embodiments may be applied to an electronic device. As shown in FIG7 , the electronic device may include a processor 31 and a memory 32. The memory 32 is used to store a program that supports the electronic device in executing the data query methods provided in the embodiments shown in FIG1 to FIG4 . The processor 31 is configured to execute the program stored in the memory 32.
[0133]
[0134] The program includes one or more computer instructions, wherein the one or more computer instructions are executed by the first processor 31 to implement the following steps: responding to a query request of the data lake, reading global statistical information of the data lake;
[0134]
[0135] Determine a query strategy corresponding to the query request according to the global statistical information; and respond to the query request according to the query strategy.
[0135]
[0136] Optionally, the processor 31 is further configured to execute all or part of the steps in the embodiments shown in FIG. 1 to FIG. 4 .
[0136]
[0137] The structure of the electronic device may further include a communication interface 33, which is used for the electronic device to communicate with other devices or communication systems.
[0137]
[0138] In addition, an embodiment of the present disclosure provides a computer storage medium for storing computer software instructions used by the above electronic device, which includes a program for executing the data query method shown in Figures 1 to 4 above.
[0138]
[0139] In addition, embodiments of the present disclosure provide a computer program product. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is enabled to implement the steps or functions of the data query method shown in Figures 1 to 4 above.
[0139]
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.
Claims
Claims 1. A data query method, comprising: In response to a query request from the data lake, read global statistical information of the data lake; determining a query strategy corresponding to the query request according to the global statistical information; Respond to the query request according to the query strategy.
2. The method according to claim 1, wherein: The method is applied to a query engine, the global statistical information of the data lake is stored in a database shared by multiple data lakes, or stored in a cache of a query device, and the query engine is deployed in the query device.
3. The method according to claim 1, further comprising: Determine global statistical information of the data lake based on statistical information of different data partitions in the data lake.
4. The method according to claim 3, wherein: The statistical information of the data partition includes the field cardinality of the data partition, and the global statistical information includes the field cardinality of the data lake; Determining global statistical information of the data lake based on statistical information of different data partitions in the data lake includes: performing statistics on characteristic values of the different data partitions stored in a first data structure to obtain characteristic values of the data lake stored in a second data structure; and converting the storage format of the characteristic values of the data lake from the second data structure to the first data structure. The field cardinality of the data lake is determined according to the characteristic values of the data lake stored in the first data structure.
5. The method according to claim 4, wherein: The first data structure includes a one-dimensional array, and the second data structure includes a two-dimensional array, wherein the length of the one-dimensional array is the same as the number of rows in the two-dimensional array. The counting of characteristic values of different data partitions stored in the first data structure to obtain characteristic values of the data lake stored in the second data structure includes: generating a two-dimensional array corresponding to the data lake based on the one-dimensional arrays corresponding to different data partitions in the data lake, wherein a target element in the second array is the number of target values in the one-dimensional array, the target element corresponds to a target number of rows and a target number of columns in the second array, the position of the target value in the one-dimensional array is the same as the target number of rows, and the target value is the same as the target number of columns.
6. The method according to claim 5, further comprising: Performing hash calculation on data in any data partition in the data lake; The hash calculation result is truncated into a first part and a second part as an index, and the columns of the two-dimensional array are The number is the same as the number of digits in the first part of the hash calculation result; the maximum number of leading zeros in the first part of the hash calculation result is written into the position pointed to by the second part of the hash calculation result in the one-dimensional array to obtain the one-dimensional array corresponding to any data partition.
7. The method according to claim 5, wherein: The converting the storage format of the characteristic values of the data lake from the second data structure to the first data structure includes: traversing different rows in the two-dimensional array in sequence according to a preset direction to obtain the first non-zero data in the different rows; and writing the non-zero data corresponding to different rows in the two-dimensional array into the one-dimensional array corresponding to the data lake in sequence according to the order in which the non-zero data corresponding to different rows in the two-dimensional array are obtained.
8. The method according to claim 5, further comprising: In response to additions and deletions of data partitions in the data lake, a two-dimensional array corresponding to the data lake is updated.
9. The method according to claim 8, wherein: The updating of the two-dimensional array corresponding to the data lake in response to additions and deletions of data partitions in the data lake includes: obtaining, in response to the addition of a data partition, a one-dimensional array corresponding to the newly added data partition and a two-dimensional array corresponding to the data lake; traversing the one-dimensional array corresponding to the newly added data partition; and incrementing by one the elements corresponding to a first reference row number and a first reference column number in the two-dimensional array corresponding to the data lake, wherein the first reference column number is the same as the first element in the one-dimensional array corresponding to the newly added data partition currently traversed, and the first reference row number and the first element have the same position in the one-dimensional array corresponding to the newly added data partition.
10. The method according to claim 8, wherein: The updating of the two-dimensional array corresponding to the data lake in response to additions and deletions of data partitions in the data lake includes: in response to deletion of a data partition, obtaining a one-dimensional array corresponding to the deleted data partition and a two-dimensional array corresponding to the data lake; traversing the one-dimensional array corresponding to the deleted data partition; and subtracting one from the elements corresponding to the second reference row number and the second reference column number in the two-dimensional array corresponding to the data lake, wherein the second reference column number is the same as the second element in the one-dimensional array corresponding to the deleted data partition currently traversed, and the second reference row number and the second element have the same position in the one-dimensional array corresponding to the deleted data partition.
11. A query engine comprising: A processing component and a query optimization component; wherein the processing component is used to respond to a query request from the data lake and read the global statistical information of the data lake; The query optimization component is configured to respond to the query request according to the query strategy determined by the query optimization component; the query optimization component is configured to determine the query strategy corresponding to the query request based on the global statistical information.
12. The query engine according to claim 11, wherein: The statistical information of the data partition includes the field cardinality of the data partition, the global statistical information includes the field cardinality of the data lake, and the field cardinality is stored in the form of a first data structure; the processing component is configured to perform statistics on the field cardinality of different data partitions stored in the first data structure to obtain the field cardinality of the data lake stored in the form of a second data structure; The storage format of the field cardinality of the data lake is converted from the second data structure to the first data structure.
13. An electronic device, comprising: A memory and a computing system; wherein the memory stores executable code, and when the executable code is executed by the computing system, the computing system executes the data query method according to any one of claims 1 to 10.
14. A non-transitory machine-readable storage medium, wherein: The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by a computing system of an electronic device, the computing system executes the data query method according to any one of claims 1 to 10.
15. A computer program product comprising a computer program or instructions, wherein: When the computer program or instruction is executed by a processor, the processor is enabled to implement the data query method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Updating statistics in distributed databases
CN104769583A
Distributed histogram computing framework using data stream sketches and samples
CN115769195A
Formulating global statistics for distributed databases
US20150169688A1