Multi-source data query analysis method, query engine, electronic equipment and storage medium
By directly executing queries in the native database and supporting custom extended syntax, the multi-source data query analysis method is solved, and efficient and flexible multi-source data query analysis is achieved.
Patent Information
- Application Number
- CN202510828778.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The traditional multi-source data query and analysis process has high resource overhead, making it difficult to cope with high concurrency and massive data scenarios, cannot meet the customization of query processes, and thus adapt to complex business rules, and is unable to compatible with standardized pipelines such as Splunk and Kusto.
It provides a multi-source data query analysis method. Through pluggable design parser, executor, standardizer, cache and formatter interfaces, query is directly executed in the native database, supporting custom extended syntax, realizing multi-query syntax compatibility, and the pluggable design interface meets business needs and adapts to complex business rules and standardized pipelines.
It reduces memory computing, reduces server resource usage, supports customization of query processes, improves flexibility and compatibility, adapts to high concurrency and massive data scenarios, and is lightweight to deploy.
Smart Images

Figure CN120336371A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data query and analysis, and more particularly to a multi-source data query and analysis method, a query engine, an electronic device, and a storage medium. Background Art
[0002] In the multi-source data analysis scenario, the data scale is huge and the sources are diverse. When traditional query engines (such as Presto, Impala) perform query and analysis on multi-source data (i.e., databases), they often first pull the data from each data source into memory, and then perform correlation calculation and analysis in memory.
[0003] The above process requires large memory computing, with high resource overhead. When facing large-scale data, it is prone to problems such as exhaustion of memory resources and decline in computing efficiency. At the same time, traditional query engines are difficult to meet the requirements of customization of modern query processes (such as integrating enterprise-specific data filtering rules and business logics) and lightweight deployment of the engine (such as embedding in third-party systems). In addition, they cannot be compared with standardized pipelines (efficient tools such as Splunk, Kusto).
[0004] In summary, the traditional multi-source data query and analysis process has technical problems such as high resource overhead, difficulty in coping with high-concurrency and massive data scenarios, inability to meet the customization of query processes, thus being unable to adapt to complex business rules, and inability to be compatible with standardized pipelines such as Splunk and Kusto. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a multi-source data query and analysis method, a query engine, an electronic device, and a storage medium to alleviate the technical problems of high resource overhead in the traditional multi-source data query and analysis process, difficulty in coping with high-concurrency and massive data scenarios, inability to meet the customization of query processes, thus being unable to adapt to complex business rules, and inability to be compatible with standardized pipelines such as Splunk and Kusto.
[0006] In a first aspect, an embodiment of the present invention provides a multi-source data query and analysis method, which is applied to a query engine used as an SDK embedding or running as an independent service. The query engine includes: a parser interface with a pluggable design, an executor interface with a pluggable design, a standardizer interface with a pluggable design, a cache interface with a pluggable design, and a formatter interface with a pluggable design. The method includes: The parser interface parses the query statement input by the user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. The parser interface for fast query syntax supports custom extended syntax; The executor interface executes the corresponding database query according to the native database query statement to obtain the data query result; The normalizer interface performs normalization processing on the data query result to obtain the normalized data query result; The cache interface caches the data query results after the standardization process; The formatter interface formats the standardized data query results output by the normalizer interface and / or all standardized data query results in the cache interface, and returns the formatted data query results to the front end for display.
[0007] Furthermore, the parser interface is called and used separately.
[0008] Furthermore, the custom extended syntax is used to express the query requirements of the front end; the custom extended syntax adopts the JSON format, and the custom extended syntax supports query result set restrictions.
[0009] Furthermore, the query result set restrictions include: If the user does not provide a result restriction parameter, the query result set restriction remains as is; If the original query information in the native database query statement does not set a result quantity limit, and the user provides the result limit parameter, then adding the query result set limit of the result limit parameter to the native database query statement; If the original query information includes a result quantity limit and the user provides the result limit parameter, the smaller of the result quantity limit and the result limit parameter is used as the query result set limit of the native database query statement.
[0010] Furthermore, the parser interface of the pipeline grammar is aligned with the Kusto grammar, and the native database query statement parsed by the pipeline grammar parser is an optimized native database query statement, and the optimized native database query statement is a multi-level SQL nested sub-query statement.
[0011] Furthermore, when the query statement is a pipeline syntax query statement, the pipeline syntax parser uses the following optimization method when parsing the pipeline syntax query statement: Parsing the pipeline syntax query statement into an abstract syntax tree, wherein the abstract syntax tree includes a plurality of pipeline tasks; Traverse the plurality of pipeline tasks from back to front; Obtain the current pipeline task and the adjacent pipeline tasks of the current pipeline task, and optimize the obtained current pipeline task and adjacent pipeline tasks using the corresponding executor interfaces; Determine whether the optimization is successful; If successful, update the current pipeline task to obtain an updated pipeline task; Determine whether the updated pipeline task is reduced; If reduced, use the updated pipeline task as the multiple pipeline tasks, and return to execute the step of traversing the multiple pipeline tasks from back to front; If unsuccessful, or not reduced, move the pointer forward; Determine whether the pointer is valid; If valid, return to execute the step of obtaining the current pipeline task and the adjacent pipeline tasks of the current pipeline task; If invalid, clear the current parsing cache result, and use the finally obtained updated pipeline task as the optimized pipeline task; Determine an optimized native database query statement according to the optimized pipeline task.
[0012] Furthermore, the number of the executor interfaces is multiple, and each executor interface corresponds to a database at the backend.
[0013] In a second aspect, an embodiment of the present invention further provides a query engine that can be embedded as an SDK or run as an independent service. The query engine includes: a pluggable parser interface, a pluggable executor interface, a pluggable normalizer interface, a pluggable cache interface, and a pluggable formatter interface; The parser interface is used to parse a query statement input by a user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. The parser interface for fast query syntax supports custom extended syntax; The executor interface is used to execute a corresponding database query according to the native database query statement to obtain a data query result; The normalizer interface is used to perform normalization processing on the data query result to obtain a normalized data query result; The cache interface is used to cache the normalized data query result; The formatter interface is used to format the data query results after the normalization process output by the normalizer interface and / or all the data query results after the normalization process in the cache interface, and return the formatted data query results to the front end for display.
[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of the above first aspects are implemented.
[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and run by the processor, the machine-executable instructions cause the processor to run the method according to any one of the above first aspects.
[0016] In an embodiment of the present invention, a multi-source data query and analysis method is provided, which is applied to a query engine used as an SDK embedded or running as an independent service. The query engine includes: a pluggable parser interface, a pluggable executor interface, a pluggable normalizer interface, a pluggable cache interface, and a pluggable formatter interface. The method includes: the parser interface parses the query statement input by the user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, a parser interface for natural language query, and the parser interface for fast query syntax supports custom extended syntax; the executor interface executes the corresponding database query according to the native database query statement to obtain a data query result; the normalizer interface performs normalization processing on the data query result to obtain a normalized data query result; the cache interface caches the normalized data query result; the formatter interface formats the normalized data query result output by the normalizer interface and / or all the normalized data query results in the cache interface, and returns the formatted data query result to the front-end for display. From the above description, it can be seen that in the multi-source data query and analysis method of the present invention, the corresponding database query is directly executed in the native database according to the native database query statement, that is, the query is directly pushed down to the data source for execution, avoiding the pulling of all data, reducing memory calculation, and reducing server resource occupancy. In addition, the parser interface includes a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, a parser interface for natural language query, and the parser interface for fast query syntax supports custom extended syntax, achieving compatibility with multiple query syntaxes and also supporting custom extended syntax, meeting the customization of the query process, adapting to complex business rules. In addition, the parser interface for pipeline syntax can dock pipeline syntax query statements and be compatible with standardized pipelines. In addition, each interface in the query engine is designed to be pluggable, and any interface can be customized according to business requirements as long as the interface semantics are met, with high flexibility. Moreover, the query engine can be used as an SDK embedded or running as an independent service, with a lightweight overall design and convenient deployment, alleviating the technical problems in the traditional multi-source data query and analysis process, such as large resource overhead, difficulty in coping with high-concurrency and massive data scenarios, inability to meet the customization of the query process, and thus unable to adapt to complex business rules, and inability to be compatible with standardized pipelines such as Splunk and Kusto. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 Flowchart of a multi-source data query and analysis method provided by an embodiment of the present invention; Figure 2 Structural schematic diagram of a query engine provided by an embodiment of the present invention; Figure 3 Interface semantic design schematic diagram of a parser interface provided by an embodiment of the present invention; Figure 4 Interface semantic design schematic diagram of an executor interface provided by an embodiment of the present invention; Figure 5 Interface semantic design schematic diagram of a normalizer interface provided by an embodiment of the present invention; Figure 6 Interface semantic design schematic diagram of a cache interface provided by an embodiment of the present invention; Figure 7 Interface semantic design schematic diagram of a formatter interface provided by an embodiment of the present invention; Figure 8 Deployment schematic diagram of a query engine provided by an embodiment of the present invention; Figure 9 Schematic diagram of the json format provided by an embodiment of the present invention; Figure 10 Schematic diagram of an optimizer process provided by an embodiment of the present invention; Figure 11 Schematic diagram of registering an executor interface to a controller provided by an embodiment of the present invention; Figure 12 Schematic diagram of an executor interface provided by an embodiment of the present invention; Figure 13 Definition schematic diagram of the logical implementation provided by an embodiment of the present invention; Figure 14 Schematic diagram of AST information extraction provided by an embodiment of the present invention; Figure 15 Schematic diagram of a structure sequence provided by an embodiment of the present invention; Figure 16 Schematic diagram of the SqlQuery structure definition provided by an embodiment of the present invention; Figure 17 Schematic diagram of the Rebuild merge structure provided by an embodiment of the present invention; Figure 18 Schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0019] Next, the technical solutions of the present invention will be described clearly and completely in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] The traditional multi-source data query and analysis process has high resource overhead, is difficult to handle high-concurrency and massive data scenarios, cannot meet the customization of the query process, and thus cannot adapt to complex business rules, and cannot be compatible with standardized pipelines such as Splunk and Kusto.
[0021] Based on this, in the multi-source data query and analysis method of the present invention, the corresponding database query is directly executed in the native database according to the native database query statement, that is, the query is directly pushed down to the data source for execution, avoiding the pulling of all data, reducing in-memory computing, and reducing the occupation of server resources. In addition, the parser interface includes a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. Among them, the parser interface for fast query syntax supports custom extended syntax, realizing the compatibility of multiple query syntaxes, and also supports custom extended syntax, meeting the customization of the query process and being able to adapt to complex business rules. In addition, the parser interface for pipeline syntax can dock pipeline syntax query statements and be compatible with standardized pipelines. In addition, each interface in the query engine is of a pluggable design, and any interface among them can be customized according to business requirements as long as the interface semantics are met, with high flexibility. Moreover, the query engine can be embedded and used as an SDK and can also run as an independent service, with a lightweight overall design and convenient deployment.
[0022] To facilitate the understanding of this embodiment, first, a multi-source data query and analysis method disclosed in the embodiments of the present invention will be introduced in detail.
[0023] Embodiment 1: According to an embodiment of the present invention, an embodiment of a multi-source data query and analysis method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.
[0024] Figure 1It is a flowchart of a multi-source data query and analysis method according to an embodiment of the present invention. As Figure 1 shown, the method includes the following steps: Step S102, the parser interface parses the query statement input by the user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. The parser interface for fast query syntax supports custom extended syntax; In an embodiment of the present invention, the above multi-source data query and analysis method can be applied to a query engine used as an SDK embedding or running as an independent service. Refer to Figure 2 , the query engine includes: a pluggable parser interface, a pluggable executor interface, a pluggable normalizer interface, a pluggable cache interface, and a pluggable formatter interface. The above interfaces can be customized according to needs, as long as the specified interface semantics are met, and the flexibility is good. Figure 2 Among them, the query type custom parameter exists when the database table is created. Different query types correspond to different custom parameters. The field metadata is used to describe the distribution of the database. For example, it includes fields, field types, etc.
[0025] The above parser interface (such as the Doris SQL converter) can mask the storage differences, generate native database query statements adapted to different databases, and push them down to each database for execution, as long as the specified interface semantics are met; for common syntax, the built-in parser interfaces include: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. The parser interface for fast query syntax supports custom extended syntax, and their corresponding query statements are custom syntax query statements, SQL syntax query statements, pipeline syntax query statements, and natural language query statements respectively.
[0026] In the parser interface, the parser interface for advanced query SQL syntax, the parser interface for pipeline syntax, and the parser interface for natural language query can directly pass their own business parameters, and the parser interface for fast query syntax can implement customized parsing optimization for the business. The interface semantics design of the parser interface is as Figure 3 shown.
[0027] Step S104, the executor interface executes the corresponding database query according to the native database query statement to obtain the data query result; Specifically, for different backend databases / large database systems, different executor interface implementations are corresponding. The interface semantics design of the executor interface is as Figure 4 shown.
[0028] Step S106, the normalizer interface performs normalization processing on the data query result to obtain the normalized data query result; Specifically, the above-mentioned normalized data query result facilitates subsequent normalization operations, such as sorting, etc., such as the commonly used json / csv. The normalization process can be customized by the normalizer interface in combination with business information, such as including other enrichment information. The interface semantics design of the normalizer interface is as Figure 5 shown.
[0029] Step S108, the cache interface caches the normalized data query result; Specifically, the present invention introduces a cache interface to uniformly optimize performance issues, especially in the paging scenario. For large data storage, batch queries are preferred first, and the results are cached in the cache interface. Paging no longer requests the library. Multiple storage connections can be implemented according to the cache interface, such as, memory / mongo / postgre sql, etc. The interface semantics design of the cache interface is as Figure 6 shown.
[0030] Step S110, the formatter interface formats the normalized data query result output by the normalizer interface and / or all the normalized data query results in the cache interface, and returns the formatted data query result to the front end for display.
[0031] Specifically, the formatting process is customized as needed. The interface semantics design of the formatter interface is as Figure 7 shown.
[0032] In an embodiment of the present invention, a multi-source data query and analysis method is provided, which is applied to a query engine used as an SDK embedding or running as an independent service. The query engine includes: a pluggable parser interface, a pluggable executor interface, a pluggable normalizer interface, a pluggable cache interface, and a pluggable formatter interface. The method includes: the parser interface parses the query statement input by the user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. The parser interface for fast query syntax supports custom extended syntax; the executor interface executes the corresponding database query according to the native database query statement to obtain a data query result; the normalizer interface performs normalization processing on the data query result to obtain a normalized data query result; the cache interface caches the normalized data query result; the formatter interface formats the normalized data query result output by the normalizer interface and / or all the normalized data query results in the cache interface, and returns the formatted data query result to the front-end for display. Through the above description, it can be seen that in the multi-source data query and analysis method of the present invention, the corresponding database query is directly executed in the native database according to the native database query statement, that is, the query is directly pushed down to the data source for execution, avoiding the pulling of all data, reducing in-memory computing, and reducing the occupation of server resources. In addition, the parser interface includes a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. Among them, the parser interface for fast query syntax supports custom extended syntax, realizing the compatibility of multiple query syntaxes, and also supporting custom extended syntax, meeting the customization of the query process, and being adaptable to complex business rules. In addition, the parser interface for pipeline syntax can dock pipeline syntax query statements and be compatible with standardized pipelines. In addition, each interface in the query engine is designed to be pluggable, and any interface among them can be customized according to business requirements, as long as the interface semantics are met, with high flexibility. And, this query engine can be used as an SDK embedding and also run as an independent service, with a lightweight overall design and convenient deployment, alleviating the technical problems of large resource overhead in the traditional multi-source data query and analysis process, being difficult to handle high-concurrency and massive data scenarios, being unable to meet the customization of the query process, and thus being unable to adapt to complex business rules, and being unable to be compatible with standardized pipelines such as Splunk and Kusto.
[0033] The above content briefly introduces the multi-source data query and analysis method of the present invention. The following will describe the specific content involved in detail.
[0034] In an alternative embodiment of the present invention, the parser interface is called and used separately.
[0035] Specifically, the overall design of the query engine is very lightweight. As Figure 8 shown, the supported deployment modes include: the parser interface is called separately; the query engine is embedded as an SDK; the query engine runs as an independent service.
[0036] In an alternative embodiment of the present invention, a custom extended syntax is used to express the query requirements of the front end; the custom extended syntax adopts the JSON format, and the custom extended syntax supports query result set limitations.
[0037] Specifically, for the queries of the customer front end, generally only condition combinations are simply passed, and business rules cannot be fixed; here, a DSL (Domain Specific Language) is customized for business requirements. By designing and implementing a specific domain language (DSL), the query requirements of the customer front end can be met. It allows users to express query conditions in a structured and flexible manner without having to directly write complex SQL or other database query codes. The DSL can use json to quickly query common functions (because it is easy to generate and parse, and can well represent hierarchical query conditions), and directly implement the built-in parser interface. A common json format is as Figure 9 shown.
[0038] In an alternative embodiment of the present invention, the query result set limitations include: (1) If the user does not provide a result limitation parameter, the query result set limitation remains unchanged; (2) If the original query information in the native database query statement does not set a result quantity limitation, and the user provides a result limitation parameter, then add the query result set limitation of the result limitation parameter to the native database query statement (that is, use the result limitation parameter as the result quantity limitation in the original query information of the native database query statement); (3) If the original query information contains a result quantity limitation, and the user provides a result limitation parameter, then use the smaller value of the result quantity limitation and the result limitation parameter as the query result set limitation of the native database query statement.
[0039] Specifically, the above query result set limit can reduce network and memory overhead. Specifically: it is allowed to pass in the result_limit parameter to specify the maximum number of results returned by the query, preventing system overload, and the following conditions need to be met: 1. If no limit is passed in, the query remains unchanged; 2. If the current query has no limit, the passed-in result_limit is added to the query by default; 3. If the current query already has a limit, compare it with result_limit and take the smaller value.
[0040] The above process is described as follows: 1. If no result_limit is passed in, the query remains unchanged This means that if the caller (user) does not provide a result_limit parameter (result limit parameter, i.e., does not specify the desired maximum number of results), then the default behavior of the system is not to limit the number of query results in any way. In other words, the query will be executed as originally and may return any number of results.
[0041] 2. If the current query has no limit, the passed-in result_limit is added to the query by default If there is no limit set in the current query (that is, the original query itself does not limit the number of returned results), and the caller provides a result_limit parameter, then the system should automatically add a limit clause to the query, with its value equal to the passed-in result_limit. This ensures that even if the original query was not designed to control the result quantity, the system can be protected from overload through the passed-in parameter.
[0042] 3. If the current query has a limit, compare it with result_limit and take the smaller value In some cases, the original query may already contain a limit clause specifying the maximum number of returned results. If the caller also provides a result_limit at this time, then a more restrictive (i.e., smaller numerical value) limit needs to be selected between these two limits. The purpose of this is to ensure that regardless of how the original query was designed, the number of records finally returned to the user will not exceed result_limit, thus further preventing problems caused by excessive data volume.
[0043] In an optional embodiment of the present invention, the parser interface of the pipeline syntax is benchmarked against the Kusto syntax, and the native database query statement parsed by the parser of the pipeline syntax is an optimized native database query statement, and the optimized native database query statement is a subquery statement with multi-level SQL nesting.
[0044] Specifically, the pipeline syntax is converted into a multi-level SQL nested subquery statement and pushed down to the data source for execution. The pipeline syntax is as follows: select src_ip, dst_ip, datatime, log_type, dst_port from qt_eventwhere activity='net_connect' | datatime>"now-1d" | dst_port in (7001, 7002, 22) and log_type = 0 | group by src_ip, dst_ip | group by src_ip | order by cnt desc The multi-level SQL nested subquery statement (i.e., the optimized native database query statement) is as follows: SELECT src_ip, count(1) AS cnt FROM (SELECT src_ip, dst_ip, count(1) AS cnt FROM (SELECT src_ip, dst_ip, datatime, log_type, dst_port FROM qt_event WHERE activity = 'net_connect' AND datatime>"2025-03-12 20:24:28.000" AND dst_port IN (7001, 7002, 22) AND log_type = 0) AS sub GROUP BY src_ip, dst_ip) AS sub GROUP BY src_ip ORDER BY cnt DESC While supporting pipeline syntax, the present invention also realizes lightweight pipeline optimization. The core lies in merging adjacent pipelines and conditions as much as possible through a multi-level merging strategy, and finally converting the pipeline into a subquery statement nested with multi-level SQL. In an optional embodiment of the present invention, when the query statement is a pipeline syntax query statement, the optimizer of the pipeline syntax, when parsing the pipeline syntax query statement, adopts the following optimization method: (1) Parse the pipeline syntax query statement into an abstract syntax tree, where the abstract syntax tree includes multiple pipeline tasks; (2) Traverse multiple pipeline tasks from back to front; (3) Obtain the current pipeline task and the adjacent pipeline task of the current pipeline task, and optimize the obtained current pipeline task and adjacent pipeline task by using the corresponding executor interface; (4) Determine whether the optimization is successful; (5) If successful, update the current pipeline task to obtain the updated pipeline task; (6) Determine whether the updated pipeline task is reduced; (7) If reduced, use the updated pipeline task as multiple pipeline tasks, and return to execute the step of traversing multiple pipeline tasks from back to front; (8) If not successful, or not reduced, move the pointer forward; (9) Determine whether the pointer is valid; (10) If valid, return to execute the step of obtaining the current pipeline task and the adjacent pipeline task of the current pipeline task; (11) If invalid, clear the current parsing cache result, and use the finally obtained updated pipeline task as the optimized pipeline task; (12) Determine the optimized native database query statement according to the optimized pipeline task.
[0045] Reference Figure 10 , the controller controls the entire optimization process, and will traverse all registered executor interfaces to merge the logic of adjacent two pipelines. This process is polling until all executor interfaces have been executed or no further optimization can be performed (all pipelines are merged together). The detailed process is as follows: (1) Input and initialization: Input the AST (abstract syntax tree) task list: The input of the optimizer process is an AST task list.
[0046] Initialize the Controller: Perform initialization operations on the controller to prepare for subsequent processing.
[0047] (2) Traversal and processing: Traverse pipeline tasks from back to front: The optimizer starts from the backend of the list and traverses the pipeline tasks in the pipeline forward.
[0048] (3)Task optimization: Obtain adjacent pipeline tasks: During the traversal, obtain the current and adjacent pipeline tasks.
[0049] Try all executor interface optimizations: Use different executor interfaces (such as LimitExecutor, other executors, WhereExecutor, GroupByExecutor, SelectExecutor, OrderExecutor, etc.) to perform tentative optimizations on the obtained tasks.
[0050] Judge whether the optimization is successful: After each executor interface performs the optimization, it is necessary to judge whether the optimization is successful this time.
[0051] (4)Update and judgment: Update pipeline tasks: If the optimization is successful, update the current pipeline tasks.
[0052] Task simplicity judgment: Judge whether the optimized task becomes simpler (for example, whether the complexity of the task is reduced). If the task simplicity does not decrease, move the pointer forward and enter the next round of optimization.
[0053] (5)Loop and end: Enter the next round of optimization: If the task simplicity decreases, continue the optimization; otherwise, move the pointer forward and enter the next round of optimization.
[0054] Whether the pointer is valid: Check whether the pointer is valid. If it is not valid, clean up the empty structure.
[0055] Return the optimized task: Finally, return the optimized pipeline tasks.
[0056] By dynamically selecting and executing the optimal executor interface, intelligent optimization of data processing tasks is realized, improving resource utilization and execution efficiency. At the same time, through the traversal method from back to front, the maximization of the optimization effect is ensured.
[0057] In the above process, the controller coordinates the entire optimization process. The strategy is to traverse the pipeline tasks from back to front and try to optimize adjacent pipeline tasks. All executor interfaces to be optimized must be registered with the Controller to take effect. As Figure 11 shown.
[0058] Executor interfaces such as Figure 12As shown, it supports multiple actuators (i.e., actuator interfaces), each for a specific optimization scenario, and all actuators implement the above Figure 12 interface.
[0059] WhereExecutor: Merges the filtering conditions of adjacent pipelines and supports the merging of complex condition trees.
[0060] GroupExecutor: Pushes down the grouping operation to the data source and processes the HAVING condition.
[0061] OrderByExecutor: Pushes down the sorting operation to the data source and verifies the existence of the sorting field in the projection.
[0062] LimitExecutor: Pushes down the limit operation to the data source and optimizes the paging query.
[0063] Logical implementation (ReBuild) The merging logic inside each actuator is controlled by ReBuild. The basic definition of ReBuild is as Figure 13 shown.
[0064] Taking the multi-level pipeline Where merge as an example, the corresponding process is as follows: AST information extraction is as Figure 14 shown. The above multi-level filtering conditions are extracted into a sequence of structures such as Figure 15 through the parser interface (Parser). The definition of the SqlQuery structure is as Figure 16 shown, and the Rebuild merge structure is as Figure 17 shown. The multi-level SqlQuery structure sequence will be merged into only one level after going through the internal process of FilterRebuild. Finally, the adjusted SqlQuery will be output.
[0065] In an optional embodiment of the present invention, the number of actuator interfaces is multiple, and each actuator interface corresponds to a backend database.
[0066] The method of the present invention has the following characteristics: an entire lightweight and extensible query engine; a multi-syntax and rule extension mechanism; a lightweight pipeline implementation algorithm that adapts to industry standards.
[0067] The method of the present invention directly sends conditions such as user query filtering and aggregation to the data source (such as a database or a big data platform) for execution, avoiding pulling a large amount of data to the local for calculation through the network and reducing in-memory calculation. It has the following advantages: Resource-efficient: Query pushdown reduces in-memory calculation and lowers server resource occupancy; High flexibility: Supports custom business rules and adapts to diverse business processes and syntax; Lightweight Compatibility: The minimalist architecture enables lightweight deployment, is compatible with industry-standard syntax, and enhances generality.
[0068] Example 2: The embodiment of the present invention also provides a query engine that can be embedded as an SDK or run as an independent service. The query engine embedded as an SDK or run as an independent service is mainly used to execute the multi-source data query and analysis method provided in the first embodiment of the present invention. The following provides a specific introduction to the query engine embedded as an SDK or run as an independent service provided in the embodiment of the present invention.
[0069] The query engine includes: a pluggable parser interface, a pluggable executor interface, a pluggable normalizer interface, a pluggable cache interface, and a pluggable formatter interface; The parser interface is used to parse the query statement input by the user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, a parser interface for natural language query, and the parser interface for fast query syntax supports custom extended syntax; The executor interface is used to execute the corresponding database query according to the native database query statement to obtain a data query result; The normalizer interface is used to perform normalization processing on the data query result to obtain a normalized data query result; The cache interface is used to cache the normalized data query result; The formatter interface is used to format the normalized data query result output by the normalizer interface and / or all the normalized data query results in the cache interface, and return the formatted data query result to the front-end for display.
[0070] In an embodiment of the present invention, a query engine that can be embedded as an SDK or run as an independent service is provided. The query engine includes: a parser interface with a pluggable design, an executor interface with a pluggable design, a normalizer interface with a pluggable design, a cache interface with a pluggable design, and a formatter interface with a pluggable design. The parser interface parses the query statement input by the user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. The parser interface for fast query syntax supports custom extended syntax. The executor interface executes the corresponding database query according to the native database query statement to obtain a data query result. The normalizer interface performs normalization processing on the data query result to obtain a normalized data query result. The cache interface caches the normalized data query result. The formatter interface formats the normalized data query result output by the normalizer interface and / or all the normalized data query results in the cache interface, and returns the formatted data query result to the front-end for display. Through the above description, it can be seen that in the query engine of the present invention, the corresponding database query is directly executed in the native database according to the native database query statement, that is, the query is directly pushed down to the data source for execution, avoiding the pulling of all data, reducing memory calculation, and reducing server resource occupancy. In addition, the parser interface includes a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language query. Among them, the parser interface for fast query syntax supports custom extended syntax, realizing the compatibility of multiple query syntaxes, and also supporting custom extended syntax, meeting the customization of the query process, being adaptable to complex business rules. In addition, the parser interface for pipeline syntax can dock pipeline syntax query statements and be compatible with standardized pipelines. In addition, each interface in the query engine has a pluggable design, and any interface can be customized according to business requirements as long as the interface semantics are met, with high flexibility. And this query engine can be embedded as an SDK or run as an independent service, with a lightweight overall design and convenient deployment, alleviating the technical problems of large resource overhead in the traditional multi-source data query and analysis process, being difficult to handle high-concurrency and massive data scenarios, being unable to meet the customization of the query process, thus being unable to adapt to complex business rules, and being unable to be compatible with standardized pipelines such as Splunk and Kusto.
[0071] Optionally, the parser interface is called and used separately.
[0072] Optionally, the custom extended syntax is used to express the query requirements of the front end; the custom extended syntax adopts the JSON format, and the custom extended syntax supports query result set limitation.
[0073] Optionally, the query result set limit includes: if the user does not provide a result limit parameter, the query result set limit remains unchanged; if the result quantity limit is not set in the information of the original query in the native database query statement and the user provides a result limit parameter, then add the query result set limit of the result limit parameter to the native database query statement; if the information of the original query includes a result quantity limit and the user provides a result limit parameter, then use the smaller value of the result quantity limit and the result limit parameter as the query result set limit of the native database query statement.
[0074] Optionally, the parser interface of the pipeline syntax is aligned with the Kusto syntax, and the native database query statement parsed by the parser of the pipeline syntax is an optimized native database query statement, and the optimized native database query statement is a subquery statement with multi-level SQL nesting.
[0075] Optionally, when the query statement is a pipeline syntax query statement, the optimizations adopted by the parser of the pipeline syntax when parsing the pipeline syntax query statement include: parsing the pipeline syntax query statement into an abstract syntax tree, where the abstract syntax tree includes multiple pipeline tasks; traversing the multiple pipeline tasks from back to front; obtaining the current pipeline task and the adjacent pipeline task of the current pipeline task, and optimizing the obtained current pipeline task and adjacent pipeline task using the corresponding executor interface; determining whether the optimization is successful; if successful, updating the current pipeline task to obtain an updated pipeline task; determining whether the updated pipeline task is reduced; if reduced, using the updated pipeline task as the multiple pipeline tasks and returning to execute the step of traversing the multiple pipeline tasks from back to front; if not successful, or not reduced, moving the pointer forward; determining whether the pointer is valid; if valid, returning to execute the step of obtaining the current pipeline task and the adjacent pipeline task of the current pipeline task; if invalid, clearing the current parsing cache result and using the finally obtained updated pipeline task as the optimized pipeline task; determining the optimized native database query statement according to the optimized pipeline task.
[0076] Optionally, the number of executor interfaces is multiple, and each executor interface corresponds to a database at the back end.
[0077] The device provided by the embodiments of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding contents in the foregoing method embodiments.
[0078] Such as Figure 18As shown, an electronic device 600 provided by an embodiment of the present application includes: a processor 601, a memory 602, and a bus. The memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs, the processor 601 communicates with the memory 602 through the bus, and the processor 601 executes the machine-readable instructions to perform the steps of the multi-source data query and analysis method as described above.
[0079] Specifically, the above-mentioned memory 602 and processor 601 can be general-purpose memory and processor, and no specific limitation is made here. When the processor 601 runs the computer program stored in the memory 602, it can execute the above-mentioned multi-source data query and analysis method.
[0080] The processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 601 or instructions in software form. The above-mentioned processor 601 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 602, and the processor 601 reads the information in the memory 602 and combines its hardware to complete the steps of the above method.
[0081] Corresponding to the above multi-source data query and analysis method, an embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the computer-executable instructions are called and run by the processor, the computer-executable instructions cause the processor to run the steps of the above multi-source data query and analysis method.
[0082] The query engine provided by the embodiments of the present application for embedding as an SDK or running as an independent service may be specific hardware on the device or software or firmware installed on the device, etc. The implementation principle and the technical effects produced by the query engine provided by the embodiments of the present application are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the query engine embodiments, reference may be made to the corresponding contents in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the foregoing-described system, query engine, and unit can all refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0083] In the embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.
[0084] For another example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of the device, method, and computer program product according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of the blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0085] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0086] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.
[0087] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the multi-source data query and analysis method described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs that can store program codes.
[0088] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0089] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, and are not intended to limit them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A multi-source data query and analysis method, characterized in that Applied to a query engine used as an embedded SDK or running as an independent service, the query engine includes: a parser interface with a pluggable design, an executor interface with a pluggable design, a normalizer interface with a pluggable design, a cache interface with a pluggable design, and a formatter interface with a pluggable design. The method includes: The parser interface parses the query statement input by the user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, and a parser interface for natural language queries. The parser interface for fast query syntax supports custom extended syntax; The executor interface executes the corresponding database query according to the native database query statement to obtain a data query result; The normalizer interface performs normalization processing on the data query result to obtain a normalized data query result; The cache interface caches the normalized data query result; The formatter interface formats the normalized data query result output by the normalizer interface and / or all the normalized data query results in the cache interface, and returns the formatted data query result to the front-end for display.
2. The method according to claim 1, wherein The parser interface is called separately.
3. The method according to claim 1, wherein The custom extended syntax is used to express the query requirements of the front end; the custom extended syntax uses the JSON format, and the custom extended syntax supports query result set restrictions.
4. The method according to claim 3, wherein The query result set restrictions include: If the user does not provide a result limit parameter, the query result set restriction remains unchanged; If the result quantity limit is not set in the information of the original query in the native database query statement and the user provides the result limit parameter, add the query result set restriction of the result limit parameter to the native database query statement; If the information of the original query contains a result quantity limit and the user provides the result limit parameter, use the smaller of the result quantity limit and the result limit parameter as the query result set restriction of the native database query statement.
5. The method according to claim 1, characterized in that, The parser interface for pipeline syntax is compatible with Kusto syntax. The native database query statement parsed by the parser for pipeline syntax is an optimized native database query statement, and the optimized native database query statement is a subquery statement with multi-level SQL nesting.
6. The method according to claim 1, wherein When the query statement is a pipeline syntax query statement, the optimization method adopted by the parser for pipeline syntax when parsing the pipeline syntax query statement includes: Parse the pipeline syntax query statement into an abstract syntax tree, where the abstract syntax tree includes multiple pipeline tasks; Traverse multiple pipeline tasks from back to front; Obtain the current pipeline task and the adjacent pipeline task of the current pipeline task, and use the corresponding executor interface to optimize the obtained current pipeline task and adjacent pipeline task; Judge whether the optimization is successful; If successful, update the current pipeline task to obtain an updated pipeline task; Determine whether the updated pipeline task is reduced; If reduced, use the updated pipeline task as the multiple pipeline tasks, and return to execute the step of traversing the multiple pipeline tasks from back to front; If not successful, or not reduced, move the pointer forward; Determine whether the pointer is valid; If valid, return to execute the step of obtaining the current pipeline task and the adjacent pipeline tasks of the current pipeline task; If invalid, clear the current parsing cache result, and use the finally obtained updated pipeline task as the optimized pipeline task; Determine an optimized native database query statement according to the optimized pipeline task.
7. The method according to claim 1, characterized in that, The number of the executor interfaces is multiple, and each executor interface corresponds to a database at the back end.
8. A query engine that is embedded as an SDK or runs as an independent service, characterized in that, The query engine includes: a parser interface with pluggable design, an executor interface with pluggable design, a normalizer interface with pluggable design, a cache interface with pluggable design, and a formatter interface with pluggable design; The parser interface is used to parse a query statement input by a user into a native database query statement adapted to the database. Among them, the parser interface includes: a parser interface for fast query syntax, a parser interface for advanced query SQL syntax, a parser interface for pipeline syntax, a parser interface for natural language query, and the parser interface for fast query syntax supports custom extended syntax; The executor interface is used to execute a corresponding database query according to the native database query statement to obtain a data query result; The normalizer interface is used to perform normalization processing on the data query result to obtain a normalized data query result; The cache interface is used to cache the normalized data query result; The formatter interface is used to format the normalized data query result output by the normalizer interface and / or all the normalized data query results in the cache interface, and return the formatted data query result to the front end for display.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7 above.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and run by the processor, the machine-executable instructions cause the processor to run the method described in any one of claims 1 to 7 above.
Citation Information
Patent Citations
Mixed query processing method and device based on big data
CN111221852A
Database query analysis method and device based on pipeline and computing equipment
CN111737284A
Search engine semantic conversion method and system, medium and electronic equipment
CN114443953A
Chart information generation method and device, vehicle and storage medium
CN118503241A
Method and system of processing plurality of database queries at database query engine
WO2024183900A1