Hybrid Query Method and System, and Storage Medium Based on Cloud Analysis Scenario
By dynamically selecting the query method of storage and computing separation or MPP architecture in the pre-computing query system, the distributed computing structure is optimized, and the efficiency and stability problems of the pre-computing query system in the ultra-high-dimensional environment are solved, and the query response of high-performance and high-concurrent search is achieved.
Patent Information
- Application Number
- CN202111062067.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-09-10
AI Technical Summary
In an ultra-high-dimensional environment, how to make the pre-computing query system most efficiently and more stablely utilize the pre-computing results, respond to customer queries the fastest, while avoiding the generation of a large amount of redundant data.
By obtaining query information, obtaining the meta information of the index based on pre-computation, and comparing with the meta information of the aggregated index, dynamically selecting the query method of storage and computing separation or MPP architecture, optimizing the distributed computing structure, building and updating the aggregated index, and reducing redundant data.
It realizes sub-second high-performance query response, supports high-concurrency dimension search, meets business needs, and ensures the stability of the query system.
Smart Images

Figure CN113918561B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a hybrid query method, a system, and a storage medium based on a cloud analysis scenario. Background Art
[0002] In the digital background, the scale of data faced by typical big data application scenarios grows exponentially. Even so, people still hope to more accurately, efficiently, conveniently, and intensively mine business value from the data. This poses high requirements for the query systems that process this data. Typical traditional distributed query computing frameworks need to occupy more memory resources, network resources, and CPU resources to meet the growing business needs. Therefore, some query systems based on the pre-computation theory have begun to receive attention, such as Apache Kylin and Apache Druid. Such pre-computation systems can utilize spare computing and storage resources to complete part of the calculations in advance and save these calculation results in a persistent storage medium. When a user query arrives, only a small amount of data reprocessing is required to answer the user's query. Therefore, such systems have very great advantages in terms of query response speed, throughput, etc. In addition, in order to maintain compatibility with business systems (including BI tools, reports, data analysis algorithms, etc.), pre-computed data query systems generally also provide SQL, or a language similar to SQL, just like general query systems.
[0003] For different queries, the expected selected aggregation indexes of the pre-computed query system are different. This provides a large choice for the distributed physical execution model: when a query selects a highly aggregated aggregation index (generally speaking, an aggregation index with a small number of data rows), such as an aggregation index aggregated by the age dimension, then the amount of data required to complete this query is very small (normally just 2 rows). However, if the query selects a low-aggregation aggregation index aggregated by the insurer + date dimension (generally speaking, with a large number of data rows), then the amount of data to be accessed is still very large.
[0004] In large-scale online multi-dimensional analysis, there are two fundamental problems in pre-computed query systems: dimensional explosion and cold start. For example, 10 dimensions will generate 1024 dimensional combinations; while 11 dimensions will generate 2048 dimensional combinations. Apache Kylin can significantly improve the average response time by cleverly selecting the aggregation indexes to be pre-computed. However, as the number of dimensions increases, for example, a typical current user tag system often starts from at least 500 dimensions. This makes any selection impractical. Especially in a distributed environment on the cloud, these two problems are amplified due to object storage. In such an ultra-high dimensional environment, how to enable the pre-computed query system to most efficiently and stably utilize the pre-computed results, respond to customer queries most quickly, and at the same time avoid generating a large amount of redundant data is the problem to be solved by the present invention.
[0005] In summary, the prior art has the following technical problems:
[0006] In an ultra-high dimensional environment, how to enable the pre-computed query system to most efficiently and stably utilize the pre-computed results, respond to customer queries most quickly, and at the same time avoid generating a large amount of redundant data. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a hybrid query method based on an analysis scenario on the cloud, including the steps of:
[0008] Obtain query information and obtain its index based on the query information;
[0009] Obtain the meta-information of the index based on pre-computation and compare it with the meta-information of the aggregation index;
[0010] Based on the result of the comparison, determine the query method corresponding to the meta-information, where the query method includes a query method of separating storage and computing or an MPP architecture.
[0011] Preferably, the obtaining of query information and obtaining its index based on the query information specifically includes:
[0012] Obtain an SQL query statement;
[0013] Obtain an SQL analyzer and analyze the SQL query statement into a syntax tree;
[0014] Extract query information as an index based on the syntax tree.
[0015] Preferably, the aggregation index specifically includes:
[0016] Obtain the data volume of historical query information;
[0017] If the data volume of the historical query information reaches a preset threshold, obtain the dimensions and metrics of the historical query information;
[0018] Construct an aggregation index based on the usage frequency of dimensions and metrics, and load the meta-information of the aggregation index into the object store;
[0019] Construct a new aggregation index based on the new data increment of dimensions and metrics, and delete the old aggregation index with a decreasing usage frequency;
[0020] Based on pre-computation, load the meta-information of the new aggregation index into the object store and update the meta-information of the aggregation index.
[0021] Preferably, the construction of a new aggregation index based on the new data increment of dimensions and metrics, and the deletion of the old aggregation index with a decreasing usage frequency specifically includes:
[0022] After receiving a request to construct a new aggregation index, determine whether to construct a new aggregation index based on user selection;
[0023] If it is determined to construct a new aggregation index, construct a new aggregation index based on the new data increment of dimensions and metrics at preset intervals. If it is determined not to construct a new aggregation index, stop;
[0024] Asynchronously delete the old aggregation index, that is, mark it as deletable and physically delete it during subsequent garbage collection.
[0025] Preferably, the obtaining of the meta-information of the index and the comparison with the meta-information of the aggregation index specifically includes:
[0026] Extract its meta-information from the index;
[0027] Compare the meta-information of the index with the meta-information of the aggregation index in the obtained object store;
[0028] When the meta-information of the index is the same as the meta-information of the aggregation index, the aggregation index is hit; otherwise, it is not hit.
[0029] Preferably, based on the comparison result, determine the query method corresponding to the meta-information. Among them, the query method includes the query method of separation of storage and computing or the MPP architecture, specifically including:
[0030] Obtain a rule library of costs, where the rule library of costs includes that when the two meta-informations are the same, the query method of separation of storage and computing is preferably selected; otherwise, the query method of the MPP architecture is preferably selected;
[0031] Obtain the comparison result and make a selection based on the rule library of costs;
[0032] Obtain the query result of separation of storage and computing or the MPP architecture.
[0033] Preferably, the rule library of costs further includes:
[0034] Push down the recognition of the query statement to the database with the MPP architecture;
[0035] Identify the aggregation and filtering parts in the query statement in the database;
[0036] Complete the aggregation in the database with the MPP architecture and return the query result.
[0037] A hybrid query system based on the cloud analysis scenario, characterized by including:
[0038] A query input module, configured to obtain query information and obtain its index based on the query information;
[0039] A pre-computation module, configured to obtain the meta-information of the index based on pre-computation and compare it with the meta-information of the aggregation index;
[0040] A query selection module, configured to select a query method of storage-computation separation or the MPP architecture according to whether the aggregation index is hit and return the query result.
[0041] Preferably, the query input module specifically includes:
[0042] Obtain the SQL query statement;
[0043] Obtain an SQL analyzer and analyze the SQL query statement into a syntax tree;
[0044] Extract the query information as an index based on the syntax tree.
[0045] Preferably, the query input module specifically includes:
[0046] Obtain the data volume of the historical query information;
[0047] If the data volume of the historical query information reaches a preset threshold, obtain the dimensions and metrics of the historical query information;
[0048] Construct an aggregation index based on the usage frequencies of the dimensions and metrics, and load the meta-information of the aggregation index into the object storage;
[0049] Construct a new aggregation index based on the new data increment of the dimensions and metrics, and delete the old aggregation index with a decreasing usage frequency;
[0050] Based on pre-computation, load the meta-information of the new aggregation index into the object storage and update the meta-information of the aggregation index.
[0051] Preferably, the query input module specifically includes:
[0052] After receiving a request to construct a new aggregation index, determine whether to construct a new aggregation index based on user selection;
[0053] If it is determined to build a new aggregation index, a new aggregation index is built based on the new data increment of dimensions and metrics at preset intervals. If it is determined not to build a new aggregation index, stop.
[0054] Asynchronously delete the old aggregation index, that is, mark it as deletable and physically delete it during subsequent garbage collection.
[0055] Preferably, the pre-computation module specifically includes:
[0056] Extract its meta-information from the index;
[0057] Compare the meta-information of the index with the meta-information of the aggregation index in the obtained object storage;
[0058] When the meta-information of the index is the same as the meta-information of the aggregation index, the aggregation index is hit; otherwise, it is not hit.
[0059] Preferably, the query selection module specifically includes:
[0060] Obtain a rule library of costs, where the rule library of costs includes that when two meta-informations are the same, the query method of storage-computation separation is preferentially selected; otherwise, the query method of the MPP architecture is preferentially selected;
[0061] Obtain a comparison result and make a selection based on the rule library of costs;
[0062] Obtain the query result of storage-computation separation or the MPP architecture.
[0063] Preferably, the query selection module specifically includes:
[0064] Push down the recognition of the query statement to the database of the MPP architecture;
[0065] Identify the aggregation and filtering parts in the query statement in the database;
[0066] Complete aggregation in the database of the MPP architecture and return the query result.
[0067] An electronic device includes a memory and a processor. The memory stores a computer program, and is characterized in that the computer program can implement any of the above methods when executed in the processor.
[0068] A storage medium stores a computer program, and is characterized in that the computer program can implement any of the above methods when executed in a processor.
[0069] The present invention classifies two common distributed computing architectures of a precomputation query system, provides an optimization strategy for a query system based on the precomputation theory, dynamically and intelligently selects the optimal distributed computing structure according to the meta-information of the precomputation results and the characteristics of the query itself, and realizes a sub-second high-performance query response. As a result, it can support higher high-concurrency dimension searches to meet business requirements, while ensuring the stability of the query system. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a flowchart of the hybrid query method based on the cloud analysis scenario of the present application;
[0071] Figure 2 It is a schematic diagram of the analysis result of the SQL query statement of the present application;
[0072] Figure 3 It is a bar schematic diagram of the test result based on User 2 of the present application;
[0073] Figure 4 It is a bar schematic diagram of the test result based on User 4 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that in the description of the present invention, unless otherwise clearly defined and limited, the term "storage medium" can be various media such as ROM, RAM, magnetic disk or optical disc that can store computer programs. The term "processor" can be a chip or circuit with data processing functions such as CPLD (Complex Programmable Logic Device), FPGA (Field-Programmable Gate Array), MCU (Microcontroller Unit), PLC (Programmable Logic Controller), and CPU (Central Processing Unit). The term "electronic device" can be any device with data processing and storage functions, and generally can include fixed terminals and mobile terminals. Fixed terminals such as desktop computers, etc. Mobile terminals such as mobile phones, PADs, and mobile robots, etc. In addition, the technical features involved in different embodiments of the present invention described hereinafter can be combined with each other as long as they do not conflict with each other.
[0075] Next, the present invention proposes some preferred embodiments to teach those skilled in the art to implement.
[0076] Example 1
[0077] This example provides a hybrid query method based on the cloud analysis scenario, as Figure 1 shown, including the steps:
[0078] S100. Obtain query information and get its index based on the query information;
[0079] S200. Obtain the meta-information of the index based on pre-computation and compare it with the meta-information of the aggregated index;
[0080] S300. Determine the query method corresponding to the meta-information based on the comparison result, where the query method includes a query method of separation of storage and computing or MPP architecture.
[0081] In a further example, the obtaining of the query information and getting its index based on the query information, as Figure 2 shown, specifically includes:
[0082] S110. Obtain the SQL query statement;
[0083] S120. Obtain an SQL analyzer and analyze the SQL query statement into a syntax tree;
[0084] S130. Extract the query information as an index based on the syntax tree.
[0085] In an even further example, the extracting of the query information as an index based on the syntax tree specifically includes:
[0086] S131. Obtain the dimensions and metrics of the query information based on the syntax tree;
[0087] S132. Compare the dimensions and metrics of the obtained index with those of the query;
[0088] S133. Select the matching index, where the matching index is a preset basic index.
[0089] In a further example, the aggregated index specifically includes:
[0090] S140. Obtain the data volume of the historical query information;
[0091] S150. If the data volume of the historical query information reaches a preset threshold, obtain the dimensions and metrics of the historical query information;
[0092] S160. Build an aggregated index based on the usage frequency of the dimensions and metrics, and load the meta-information of the aggregated index into the object storage;
[0093] S170. Build a new aggregation index based on the new data increment of dimensions and metrics, and delete the old aggregation index with a decreasing usage frequency;
[0094] S180. Load the meta - information of the new aggregation index into the object storage based on pre - calculation and update the meta - information of the aggregation index.
[0095] In a further embodiment, building an aggregation index based on the usage frequency of dimensions and metrics specifically includes:
[0096] S151. Analyze user behavior based on an intelligent optimization system;
[0097] S152. Determine the user's query habits based on the analysis results, including but not limited to the dimension types selected by the user and the metric ranges selected by the user;
[0098] S153. Build an aggregation index based on the user's query habits.
[0099] In a still further embodiment, building a new aggregation index based on the new data increment of dimensions and metrics, and deleting the old aggregation index with a decreasing usage frequency specifically includes:
[0100] S161. After receiving a request to build a new aggregation index, determine whether to build a new aggregation index based on user selection;
[0101] S162. If it is determined to build a new aggregation index, build a new aggregation index based on the new data increment of dimensions and metrics at preset intervals; if it is determined not to build a new aggregation index, stop;
[0102] S163. Asynchronously delete the old aggregation index, that is, mark it as deletable and physically delete it during subsequent garbage collection.
[0103] In a further embodiment, obtaining the meta - information of the index and comparing it with the meta - information of the aggregation index specifically includes:
[0104] S210. Extract its meta - information from the index;
[0105] S220. Compare the meta - information of the index with the meta - information of the aggregation index in the obtained object storage;
[0106] S230. When the meta - information of the index is the same as the meta - information of the aggregation index, the aggregation index is hit; otherwise, it is not hit.
[0107] In a further embodiment, determining the query method corresponding to the meta - information based on the comparison result, where the query method includes a query method of separation of storage and computing or an MPP architecture, specifically includes:
[0108] S310. Obtain a cost rule library, where the cost rule library includes that when two meta-informations are the same, the query method of storage-computation separation is preferentially selected; otherwise, the query method of the MPP architecture is preferentially selected.
[0109] S320. Obtain a comparison result and make a selection based on the cost rule library.
[0110] S330. Obtain the query result of storage-computation separation or the MPP architecture.
[0111] In a further embodiment, the cost rule library further includes:
[0112] S321. Push down the recognition of the query statement to the database of the MPP architecture.
[0113] S322. Identify the aggregation and filtering parts in the query statement in the database.
[0114] S323. Complete aggregation in the database of the MPP architecture and return the query result.
[0115] From the above description, it can be seen that the present invention achieves the following technical effects:
[0116] 1. By classifying the two common distributed computing architectures in the pre-computation query system, the technical effect of providing an optimization strategy for the query system based on the pre-computation theory is achieved.
[0117] 2. By the meta-information of the pre-computation result and the characteristics of the query itself, the technical effect of dynamically and intelligently selecting the optimal distributed computing structure is achieved.
[0118] 3. By implementing a sub-second high-performance query response, the technical effects of supporting a higher high-concurrency dimension search to meet business requirements and ensuring the stability of the query system are achieved.
[0119] Embodiment 2
[0120] This embodiment provides a hybrid query system based on the cloud analysis scenario, which is characterized by including:
[0121] A query input module, configured to obtain query information and obtain its index based on the query information.
[0122] A pre-computation module, configured to obtain the meta-information of the index based on pre-computation and compare it with the meta-information of the aggregated index.
[0123] A query selection module, configured to select the query method of storage-computation separation or the MPP architecture according to whether the aggregated index is hit and return the query result.
[0124] In a further embodiment, the query input module specifically includes:
[0125] Obtain an SQL query statement;
[0126] Obtain an SQL analyzer and analyze the SQL query statement into a syntax tree;
[0127] Extract query information based on the syntax tree as an index.
[0128] In a further embodiment, the query input module specifically includes:
[0129] Obtain the data volume of historical query information;
[0130] If the data volume of historical query information reaches a preset threshold, obtain the dimensions and metrics of the historical query information;
[0131] Construct an aggregation index based on the usage frequencies of dimensions and metrics, and load the meta-information of the aggregation index into an object store;
[0132] Construct a new aggregation index based on the new data increment of dimensions and metrics, and delete the old aggregation index with a decreasing usage frequency;
[0133] Based on pre-computation, load the meta-information of the new aggregation index into the object store and update the meta-information of the aggregation index.
[0134] In a still further embodiment, the query input module specifically includes:
[0135] After receiving a request to construct a new aggregation index, determine whether to construct a new aggregation index based on user selection;
[0136] If it is determined to construct a new aggregation index, construct a new aggregation index based on the new data increment of dimensions and metrics at preset intervals, and if it is determined not to construct a new aggregation index, stop;
[0137] Asynchronously delete the old aggregation index, that is, mark it as deletable and physically delete it during subsequent garbage collection.
[0138] In a further embodiment, the pre-computation module specifically includes:
[0139] Extract its meta-information from the index;
[0140] Compare the meta-information of the index with the meta-information of the aggregation index in the obtained object store;
[0141] When the meta-information of the index is the same as the meta-information of the aggregation index, the aggregation index is hit, otherwise it is not hit.
[0142] In a further embodiment, the query selection module specifically includes:
[0143] A rule library for obtaining costs, where the rule library for costs includes that when two meta-informations are the same, the query method of storage-computation separation is preferentially selected; otherwise, the query method of the MPP architecture is preferentially selected.
[0144] Obtain the comparison result and make a selection based on the rule library of costs.
[0145] Obtain the query result of storage-computation separation or the MPP architecture.
[0146] In a further embodiment, the query selection module specifically includes:
[0147] Push down the recognition of the query statement to the database of the MPP architecture.
[0148] Identify the aggregation and filtering parts in the query statement in the database.
[0149] Complete aggregation in the database of the MPP architecture and return the query result.
[0150] Embodiment III
[0151] Based on a hybrid query method for cloud-based analysis scenarios provided in this embodiment, through the steps:
[0152] S100. Obtain query information and obtain its index based on the query information;
[0153] In this embodiment, the obtained query information is to calculate the total sum of the policy amounts (sum(amount)) of insurance salespersons (seller_id) on a certain day (date).
[0154] S200. Compare the meta-information obtained by pre-calculating the index with the meta-information of the aggregation index;
[0155] Since the number of salespersons may be large, when there is no query history, initially, an aggregation index with dimensions (seller_id, date) and a metric of the total sum of the policy amounts (sum(amount)) will not be generated, so this query statement will not hit the aggregation index.
[0156] S300. Determine the query method corresponding to the meta-information based on the comparison result, where the query method includes the query method of storage-computation separation or the MPP architecture.
[0157] This query will hit the basic index. If only the data is read from the MPP and aggregated at the data end, SQL1 will involve a large amount of data scanning.
[0158] In a further embodiment, the obtaining of query information and obtaining its index based on the query information specifically includes:
[0159] S110. Obtain the SQL query statement;
[0160] Analyze the following query statement: SQL1 analyzes the total transaction amount of the salesperson with the ID of 10003 on January 1st: select sum(amount) from transactions where date = ’1.1’ and seller_id = ‘10003’.
[0161] S120. Obtain an SQL analyzer, and analyze the SQL query statement into a syntax tree, as Figure 2 shown;
[0162] S130. Obtain query information based on the syntax tree as an index.
[0163] In a further embodiment, the rule base of the cost further includes:
[0164] S321. Push down the recognition of the query statement to the database of the MPP architecture;
[0165] S322. Identify the aggregation and filtering parts in the query statement in the database;
[0166] S323. Complete the aggregation in the database of the MPP architecture and return the query result.
[0167] Due to the influence of the rule base, we identify that the aggregation and filtering parts in SQL1 can be pushed down to the MPP database, so that the aggregation is completed in the MPP database, and only one piece of data is returned, greatly reducing the data to be transmitted, thereby improving the performance.
[0168] In a further embodiment, the aggregation index specifically includes:
[0169] S140. Obtain the data volume of the historical query information;
[0170] S150. If the data volume of the historical query information reaches a preset threshold, obtain the dimensions and metrics of the historical query information;
[0171] S160. Construct an aggregation index based on the usage frequencies of the dimensions and metrics, and load the meta-information of the aggregation index into the object storage;
[0172] S170. Construct a new aggregation index based on the new data increment of the dimensions and metrics, and delete the old aggregation index with a decreasing usage frequency;
[0173] S180. Based on pre-computation, load the meta-information of the new aggregation index into the object storage and update the meta-information of the aggregation index.
[0174] Over time, if such queries are very frequent (usually there is a threshold, such as 100 such queries per day), the system will consider that building an aggregated index in advance for queries of this pattern can improve the overall performance. Then, after the pre-computation is completed, when the SQL is executed again, it will be routed to the storage-computation separation system.
[0175] In a further embodiment, building a new aggregated index based on the new data increment of dimensions and metrics, and deleting the old aggregated index with a decreasing usage frequency specifically includes:
[0176] S161. After receiving a request to build a new aggregated index, determine whether to build a new aggregated index based on user selection;
[0177] S162. If it is determined to build a new aggregated index, build a new aggregated index based on the new data increment of dimensions and metrics at preset intervals. If it is determined not to build a new aggregated index, stop;
[0178] S163. Asynchronously delete the old aggregated index, that is, mark it as deletable and physically delete it during subsequent garbage collection.
[0179] In a further embodiment, obtaining the meta-information of the index and comparing it with the meta-information of the aggregated index specifically includes:
[0180] S210. Extract its meta-information from the index;
[0181] S220. Compare the meta-information of the index with the meta-information of the aggregated index in the obtained object storage;
[0182] S230. When the meta-information of the index is the same as the meta-information of the aggregated index, the aggregated index is hit; otherwise, it is not hit.
[0183] In a further embodiment, determining the query method corresponding to the meta-information based on the comparison result, where the query method includes a query method of storage-computation separation or MPP architecture specifically includes:
[0184] S310. Obtain a cost rule library, where the cost rule library includes that when the two meta-informations are the same, the query method of storage-computation separation is preferably selected; otherwise, the query method of MPP architecture is preferably selected;
[0185] S320. Obtain the comparison result and make a selection based on the cost rule library;
[0186] S330. Obtain the query result of storage-computation separation or MPP architecture.
[0187] When the optimizer rule discovers through pattern matching that the aggregation operation is above the table scan, the aggregation operation can be pushed into the table scan to reduce the data transferred from the MPP engine back to the computing engine. Depending on the different SQLs, the data reduction can reach the GB level.
[0188] Example 4
[0189] In this example, a stress test was conducted based on the dataset of User 2, and the test results are as Figure 3 shown.
[0190] Here, Kyligence refers to the product without using this technology, and Kyligence with Tiered Storage refers to the latest product using this technology. The fixed queries here refer to the queries that can utilize the aggregation index, and it can be seen that there is no improvement. The Ad-hoc queries here refer to the queries that cannot be accelerated by the aggregation index. After using this technology, it transparently utilizes MPP acceleration, and under the concurrent stress test of two users, the performance has increased by 3 times.
[0191] Example 5
[0192] In this example, a stress test was conducted based on the dataset of User 4, and the test results are as Figure 4 shown.
[0193] Here, Kyligence refers to the product without using this technology, and Kyligence with Tiered Storage refers to the latest product using this technology. The fixed queries here refer to the queries that can utilize the aggregation index, and it can be seen that there is no improvement. The Ad-hoc queries here refer to the queries that cannot be accelerated by the aggregation index. After using this technology, it transparently utilizes MPP acceleration, and under the concurrent stress test of two users, the performance has also increased by nearly 2 times.
[0194] Example 6
[0195] An embodiment of the present invention further includes an electronic device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed in the processor, it is used to implement the above-mentioned hybrid query method based on the cloud analysis scenario. The method includes:
[0196] S100. Obtain query information and obtain its index based on the query information;
[0197] S200. Obtain the meta-information of the index based on pre-computation and compare it with the meta-information of the aggregation index;
[0198] S300. Determine the query method corresponding to the meta-information based on the comparison result, where the query method includes a query method of separation of storage and computing or an MPP architecture.
[0199] Example 7
[0200] In this embodiment, the present invention further provides a readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the above-mentioned hybrid query method based on the cloud analysis scenario. The method includes:
[0201] S100. Obtain query information and obtain its index based on the query information;
[0202] S200. Obtain the meta-information of the index based on pre-computation and compare it with the meta-information of the aggregated index;
[0203] S300. Determine the query method corresponding to the meta-information based on the comparison result, where the query method includes a query method of storage-computation separation or an MPP architecture.
[0204] Among them, the readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0205] The present invention further provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of the device can read the execution instructions from the readable storage medium, and at least one processor executes the execution instructions to enable the device to implement the methods provided by the above various embodiments.
[0206] In the above embodiments of the terminal or the server, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in conjunction with the present invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0207] It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0208] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing system. They can be concentrated on a single computing system or distributed on a network composed of multiple computing systems. Optionally, they can be implemented by program codes executable by the computing system, so that they can be stored in the storage system and executed by the computing system, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. Thus, the present invention is not limited to any specific combination of hardware and software.
[0209] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A hybrid query method based on a cloud analysis scenario, characterized in that, It includes: Obtain query information and get its index based on the query information; Obtain the meta-information of the index based on pre-computation and compare it with the meta-information of the aggregated index; The aggregated index specifically includes: Obtain the data volume of historical query information; If the data volume of historical query information reaches a preset threshold, obtain the dimensions and metrics of the historical query information; Construct an aggregated index based on the usage frequencies of dimensions and metrics, and load the meta-information of the aggregated index into object storage; Construct a new aggregated index based on the new data increment of dimensions and metrics, and delete the old aggregated index with a decreasing usage frequency; Based on pre-computation, load the meta-information of the new aggregated index into object storage and update the meta-information of the aggregated index; Based on the comparison result, determine the query method corresponding to the meta-information. Among them, the query method includes the query method of storage-computation separation or the MPP architecture, including: Obtain the rule library of costs, where the rule library of costs includes that when two meta-informations are the same, preferentially select the query method of storage-computation separation, otherwise preferentially select the query method of the MPP architecture; Obtain the comparison result and make a selection based on the rule library of costs; Obtain the query result of storage-computation separation or the MPP architecture.
2. The method according to claim 1, characterized in that, The obtaining of query information and getting its index based on the query information specifically includes: Obtain the SQL query statement; Obtain an SQL analyzer and analyze the SQL query statement into a syntax tree; Extract the query information as an index based on the syntax tree.
3. The method according to claim 1, wherein The constructing of a new aggregated index based on the new data increment of dimensions and metrics and deleting the old aggregated index with a decreasing usage frequency specifically includes: After receiving a request to construct a new aggregated index, determine whether to construct a new aggregated index based on user selection; If it is determined to construct a new aggregated index, construct a new aggregated index based on the new data increment of dimensions and metrics every preset time. If it is determined not to construct a new aggregated index, stop; Asynchronously delete the old aggregated index, that is, mark it as deletable and physically delete it during subsequent garbage collection.
4. The method according to claim 1, wherein The obtaining of the meta-information of the index based on pre-computation and comparing it with the meta-information of the aggregated index specifically includes: Extract its meta-information from the index; Compare the meta-information of the index with the meta-information of the aggregated index in the obtained object storage; When the meta-information of the index is the same as the meta-information of the aggregated index, the aggregated index is hit, otherwise it is not hit.
5. The method according to claim 4, characterized in that, The rule library of costs further includes: Push down the recognition of the query statement to the database of the MPP architecture; Identify the aggregation and filtering parts in the query statement in the database; Complete aggregation in the database of the MPP architecture and return the query result.
6. A system for hybrid query based on cloud analysis scenarios, characterized in that, It includes: A query input module for obtaining query information and getting its index based on the query information; A pre-computation module for obtaining the meta-information of the index based on pre-computation and comparing it with the meta-information of the aggregated index; The aggregated index specifically includes: Obtain the data volume of historical query information; If the data volume of historical query information reaches a preset threshold, obtain the dimensions and metrics of the historical query information; Construct an aggregated index based on the usage frequencies of dimensions and metrics, and load the meta-information of the aggregated index into object storage; Construct new aggregation indexes based on new data increments of dimensions and metrics, and delete old aggregation indexes with decreasing usage frequency; Load the meta-information of the new aggregation index into the object storage based on pre-computation and update the meta-information of the aggregation index; A query selection module, configured to determine a query method corresponding to the meta-information based on the comparison result, wherein the query method includes a query method of separation of storage and computing or an MPP architecture; The query selection module is further configured to: Obtain a rule base of costs, where the rule base of costs includes that when two meta-informations are the same, the query method of separation of storage and computing is preferentially selected, otherwise the query method of the MPP architecture is preferentially selected; Obtain the comparison result and make a selection based on the rule base of costs; Obtain the query result of separation of storage and computing or the MPP architecture.
7. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, The computer program, when executed in the processor, can implement any one of the methods in claims 1-5.
8. A storage medium stores a computer program, characterized in that, The computer program, when executed in the processor, can implement any one of the methods in claims 1-5.
Citation Information
Patent Citations
Query optimization
US20210034616A1