Data retrieval method and computing equipment
By segmenting and generating target search statements, combined with multi-level connection methods, the problem of low cross-data source search efficiency in the existing technology is solved, and efficient unified search across data sources is achieved.
Patent Information
- Application Number
- CN202510337633.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
Existing data retrieval systems usually only search for a single data source, which cannot meet the needs of cross-data source retrieval and are inefficient in retrieval.
By obtaining search statements, segmenting and determining keywords, generating target search statements based on the parameter information of each target data source, and sending them to the corresponding data source for searching, multi-level waterfall or feedback connections are used to conduct cross-data source joint query to realize cross-data source retrieval.
It realizes searching through unified search statements in multiple data sources without manually switching the system, improving data retrieval efficiency.
Smart Images

Figure CN120336348A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data retrieval, and in particular, to a data retrieval method and a computing device. Background Art
[0002] Data retrieval refers to the process of searching for and extracting data from a database or other data storage systems. In today's information age, data retrieval has become an indispensable and important tool in many fields such as people's access to knowledge, research, and business activities.
[0003] With the rapid development of the Internet and the wide application of various information systems, data sources have become increasingly rich and diverse, covering various forms such as structured databases, unstructured texts, multimedia resources, and distributed storage systems. Facing such a vast and scattered data source, currently, most existing data retrieval systems usually only retrieve for a single data source. This single-data-source retrieval method cannot meet the user's need for cross-data-source retrieval and has a low retrieval efficiency.
[0004] Therefore, how to achieve cross-data-source retrieval is a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention
[0005] The embodiments of this application provide a data retrieval method and a computing device, which can achieve cross-data-source retrieval and improve the efficiency of data retrieval.
[0006] In a first aspect, the embodiments of this application provide a data retrieval method, which is applied to a computing device. The computing device is connected to multiple target data sources. The method includes:
[0007] Obtain a retrieval statement, split the retrieval statement to determine the keywords of the retrieval statement, then generate target retrieval statements respectively corresponding to each target data source according to the keywords and the data source parameter information respectively corresponding to each target data source, and finally the target retrieval statements can be sent to the corresponding target data sources respectively, so that the target data sources use the corresponding target retrieval statements for retrieval, and determine the target retrieval result according to the retrieval results of each target data source. In this way, it allows users to perform retrieval in multiple data sources simultaneously through a unified retrieval statement, without manually switching different systems for retrieval, and can achieve cross-data-source retrieval, improving the efficiency of data retrieval.
[0008] In some possible implementations, this application can perform cross-data source federated queries using a multi-level waterfall join or a multi-level feedback join. In a multi-level waterfall join (Waterfall Join), the query passes conditions step by step, and the result of each query serves as the filtering condition for the next query. The final result includes the result of the last query, that is, execute the first-level query (execute the first query to obtain the result set) and then pass the condition (extract the value of the specified field from the result set of the first-level query as the filtering condition for the second-level query). Here, the specified field can also be called the target field, which is a field used to locate, filter, or associate data in the data source. For example, the value of a certain ID field can be extracted from the result of the first-level query as the target field. The filtering condition means using the value of the target field as the filtering condition to limit the data range returned by the query, and then perform subsequent queries (repeat the above process, passing the result of each level of query to the next level of query until the last query) to obtain the final result (the result set of the last query is used as the final result).
[0009] In a multi-level feedback join (Feedback Join), the result of each query is associated and concatenated with the result of the previous query to gradually build the final result set. Each query is executed based on the result of the previous query and the results are merged. That is, execute the first-level query (execute the first query to obtain the result set), then pass the condition (extract the value of the specified field from the result set of the first-level query as the filtering condition for the second-level query), and then perform the association and concatenation (the result of the second-level query will be associated and concatenated with the result of the first-level query through the specified key value (which can also be called the target key value, such as a primary key or a foreign key, for example, user_id)), perform subsequent queries (repeat the above process, associating and concatenating the result of each level of query with the result of the previous level of query until the last query), and obtain the final result: (perform the final concatenation on all the gradually associated and concatenated results, and the obtained result is used as the final result).
[0010] In some possible implementations, split the retrieval statement to determine the keywords of the retrieval statement. Specifically, it can be:
[0011] The retrieval statement can be split according to the data source index keyword and the target data syntax command keyword to obtain the keywords of the sub-retrieval statement corresponding to each data source.
[0012] Generate the target retrieval statements corresponding to each target data source. Specifically, it can be:
[0013] According to the keywords of the sub-search statement, the operation object of the keywords, and the conditions after the keywords, each sub-search statement is converted into a structured statement according to the grammar rules corresponding to the data source, and then the structured statement is converted into a target search statement corresponding to the target data source.
[0014] In some possible implementations, the structured statement may include at least one of a data source index, a search condition, a filtering condition, a statistical function, an aggregation function, a non-aggregation function, a merge command, a sorting command, a command to retrieve the first N data records, and a multi-data source fusion command. Among them, the data source index is used to represent the source of the data in the sub-search statement, the search condition is used to filter out the records that meet the conditions in the data source of the sub-search statement, the filtering condition is used to remove duplicate records and filter specific fields in the sub-search statement, the statistical function is used to perform statistical analysis on the data in the sub-search statement, the aggregation function is used to merge multiple rows of data in the sub-search statement into a single row of data, the non-aggregation function is used to operate on a single row of data in the sub-search statement, the merge command is used to merge multiple query results into a single result set, the sorting command is used to sort the query results, the command to retrieve the first N data records is used to limit the number of query results, and the multi-data source fusion command is used to retrieve data from multiple data sources and merge the results into a single result set.
[0015] In some possible implementations, converting the structured statement into a target search statement corresponding to the target data source may specifically be:
[0016] Extract the placeholder or data source index in the structured statement, extract the required parameter values from the corresponding data source parameter information, replace the placeholder or data source index with the extracted parameter values, and convert the replaced statement into a target search statement corresponding to the data source.
[0017] In some possible implementations, sending the target search statements to the corresponding target data sources respectively may specifically be:
[0018] Construct an execution queue according to the target search statements, where the execution queue is used to determine the execution order of each target search statement, and then send the target search statements to the corresponding target data sources respectively from the execution queue.
[0019] In some possible implementations, the present application may also map the storage levels of different target data sources to databases and tables, obtain the mapping relationship corresponding to each data source, where the mapping relationship is used to map the respective corresponding storage levels in different heterogeneous data sources to a unified level, and store the mapping relationship.
[0020] In a second aspect, an embodiment of the present application provides a computing device, including: a memory and a processor;
[0021] The memory is coupled to the processor;
[0022] The memory stores program instructions that, when executed by the processor, cause the computing device to perform the method described in the first aspect.
[0023] As can be seen from the above technical solutions, the present application has the following beneficial effects:
[0024] The present application can access multiple target data sources, and can convert the obtained retrieval statements into target retrieval statements respectively corresponding to each target data source, and then send them to the corresponding target data sources for retrieval. This method allows users to retrieve in multiple data sources simultaneously through a unified retrieval statement, without the need to manually switch different systems for retrieval, and can achieve cross-data-source retrieval, improving the efficiency of data retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of a data retrieval method provided by an embodiment of the present application;
[0026] Figure 2 It is a flowchart of another data retrieval method provided by an embodiment of the present application;
[0027] Figure 3 It is a hardware structure of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0029] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0030] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe a specific order of the target objects.
[0031] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0032] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of systems refers to two or more systems.
[0033] The following introduces relevant technical terms and the application scenarios of the solutions of the present application.
[0034] Data retrieval: Data retrieval refers to the process of searching for and extracting data from a database or other data storage systems. This process usually involves using a query language to specify the conditions and formats of the required data. An effective data retrieval system should be able to quickly and accurately return the information required by the user, while supporting complex query logic.
[0035] Multiple data sources: Multiple data sources refer to using data from different storage systems in the same system or application. These data sources can include relational databases, NoSQL databases, file systems, cloud storage services, etc. Integrating multiple data sources requires solving problems such as data format standardization and data consistency to ensure the effective use of data throughout the system.
[0036] Unified language: Unified language refers to using a consistent programming or query language between different data processing components or systems. This language is designed to simplify the data processing process, reduce the learning cost, and improve the development efficiency. A unified language can help developers switch more easily between different data sources and technology stacks without having to learn multiple different languages or interfaces. For example, SQL is a widely used unified language that can be used in multiple database management systems.
[0037] Associative query: In a database management system, an operation of combining data from multiple tables to generate a new result set.
[0038] The data retrieval method provided by this application can be applied to computing devices such as servers or terminals. The computing device accesses multiple target data sources. The method of this application may include obtaining a retrieval statement, splitting the retrieval statement, determining the keywords of the retrieval statement, and then converting the retrieval statement into target retrieval statements respectively corresponding to each target data source according to the keywords and the data source parameter information respectively corresponding to each target data source. Furthermore, the target retrieval statements can be respectively sent to the corresponding target data sources so that the target data sources use the corresponding target retrieval statements for retrieval. Finally, the target retrieval result can be determined according to the retrieval results of each of the target data sources. In this way, the computing device of this application accesses multiple target data sources and can convert the retrieval statement into target retrieval statements respectively corresponding to each target data source, and thus send them to the corresponding target data sources for retrieval. This method allows users to retrieve data in multiple data sources simultaneously through a unified retrieval statement, without manually switching different systems for retrieval, and can achieve cross-data source retrieval, improving the efficiency of data retrieval.
[0039] It should be noted that the solution of this application can be applied to computing devices in a computer cluster. The system architecture and application scenarios described in the embodiments of this application are for more clearly explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0040] For the sake of easy understanding, the data retrieval method provided by this application is introduced exemplarily below with reference to the accompanying drawings, which can be referred to as Embodiment 1 here. Taking the execution of the method in Embodiment 1 in a computing device as an example for introduction. The data retrieval method provided in Embodiment 1 of this application is introduced in detail below. Specifically, as Figure 1 shown, the data retrieval method includes:
[0041] S101. Obtain the input retrieval statement.
[0042] The computing device obtains the retrieval statement input by the user. Here, the retrieval statement refers to a sequence of instructions or commands for finding, filtering, processing, or extracting information from a data source. It can be written according to certain syntax and logical rules to complete specific data retrieval tasks.
[0043] Exemplarily, the keywords of the retrieval statement may include data source index keywords (data source index keywords are keywords used to specify the retrieval scope or data source. For example, in an SQL query, the data source can be specified through table names, field names, etc. In a search engine, the retrieval scope can be narrowed down through specific site restrictions or file type restrictions. In file retrieval, the retrieval scope can be specified through paths or file extensions, etc.), target data syntax command keywords (target data syntax command keywords are keywords used to describe the format, structure, or syntax rules of retrieving target data), etc.
[0044] For example, the retrieved input retrieval statement is "datafilter mysql db=TEST table=test1·search*", where datafilter is a data source index keyword, and search is a target data syntax command keyword. An operation object or operation condition can be appended after search.
[0045] The data source index syntax rule of this retrieval statement (the data source index syntax rule is a format specification used to clearly specify the data source of the retrieved data. Through a fixed syntax structure, it can quickly locate the database, table, or other data storage locations where the target data is located) can be: "datafilter is a fixed keyword indicating that this is an operation for data filtering and is also used to specify the type of data source. db represents the name of the database to be retrieved, and table represents the name of the table to be retrieved.
[0046] Specifically, the target data syntax command keywords can be: retrieval conditions, filtering conditions, statistical functions, aggregation functions, non-aggregation functions, merge commands, sorting commands (which can also be called sorting conditions, etc.), commands for fetching the first N data records, and multi-data source fusion commands, etc. For example, they can be search (search command and filtering condition), limit (command for fetching the first N data records), sort (sorting command), deduplication (deduplicating data according to the specified field list), and join (multi-data source fusion command), etc. Among them, the data source index is used to represent the source of the data in the sub-retrieval statement. The retrieval condition is used to filter out the records that meet the conditions in the data source of the sub-retrieval statement. The filtering condition is used to remove duplicate records in the sub-retrieval statement and filter specific fields. The statistical function is used to perform statistical analysis on the data in the sub-retrieval statement. The aggregation function is used to merge multiple rows of data in the sub-retrieval statement into a single row of data. The non-aggregation function is used to operate on a single row of data in the sub-retrieval statement. The merge command is used to merge multiple query results into one result set. The sorting command is used to sort the query results. The command for fetching the first N data records is used to limit the number of query results. The multi-data source fusion command is used to retrieve data from multiple data sources and merge the results into one result set.
[0047] For example, the search command can filter data that meets specified conditions through logical expressions, and perform primary or secondary or more filtering on the data. It can be the first command after the data source index in the retrieval statement, or the first command in the retrieval statement (the search command is equivalent to the select command in SQL). If the right value (here it is the value compared with the field (or column). That is, it is the conditional value used to filter data, such as test1·search* in the above retrieval statement) is a string and the string contains an asterisk (*), then the asterisk can be treated as a wildcard. It should be noted that here it is different from using.* to represent wildcards in regular expressions. In search, * is used to represent wildcards, and the dot [.] is only used as an ordinary character in search. For example, x = abc* or x = "abc*" means that the value of x starts with abc and is followed by zero or more arbitrary characters, which is equivalent to prefix matching. x = *abc is equivalent to suffix matching, which means that the value of x ends with abc and can be any character in front. x = *abc* is equivalent to substring wildcarding, which means that the value of x contains the substring abc and can be any character before and after. For prefix matching, suffix matching, and substring wildcarding, the syntax can be optimized for execution efficiency. For example, for prefix matching, the present application can use an index to quickly locate records starting with a specific prefix. For suffix matching, the present application can be optimized through data structures such as inverted indexes or suffix trees. For substring wildcarding, the present application can be optimized through string search algorithms (such as the KMP algorithm, Boyer-Moore algorithm, etc.).
[0048] If the right value string contains an asterisk in the middle, regular matching can be used. For example, x = a*c means that the value of x contains a and c, and there can be any characters in the middle.
[0049] The limit command is used to represent the maximum number of records returned from the data source. When the maximum number of records to be returned is not specified through the limit command in the retrieval statement, the default maximum number of records returned is 10,000. The syntax format is limit <maximum number of records>, and the maximum number of records must be an integer greater than 0 and less than 1 million. For a join executed by pushing down to Presto (a SQL query engine), the head command in the query can set the "maximum number of records" to a maximum of 10 million to allow a join operation on a dataset of tens of millions in Presto (the join operation is used to connect rows from different tables or data sources according to specified conditions to form a new result set). The query result set of each data source is defaultly limited to a maximum of 10,000. Through the limit command, it is possible to specify a maximum of 1 million data to be returned from the data source to the local computing device.
[0050] S102. Determine the keywords of the retrieval statement.
[0051] The computing device can split the retrieval statement and determine the keywords (also called key words) of the retrieval statement.
[0052] The computing device can split the retrieval statement into multiple sub-retrieval statements.
[0053] In some possible implementations, in a retrieval statement, the output value of the first sub-retrieval statement after splitting can be used as the input value of the second sub-retrieval statement after splitting, and the execution is performed according to the split result.
[0054] The splitting rule for splitting the retrieval statement can be "data source index keyword·syntax command keyword 1 for target data·syntax command keyword 2 for target data·syntax command keyword 3 for target data·...". Multiple syntax command operations can be combined by using
·
[0055] Exemplarily, for example, the retrieval statement is: "datafilter mysql db=TEST table=test1·search name=test_app field timefield,uid,ip,id_test·dedup host_name limit 50000 rename host_name as names·join type=feedback names[datafilter mongo db=mongo_test table log_http·search type=ip_address·stats count by names]". Splitting it according to the splitting rule, it can be split into the first sub-retrieval statement "datafilter mysql db=TEST table=test1·search name=test_app field timefield,uid,ip,id_test·dedup host_name limit 50000 rename host_name as names" and the second sub-retrieval statement "datafilter mongo db=mongo_test table=log_http·search type=ip_address·stats count by names". One data source can correspond to one sub-retrieval statement.
[0056] Among them, the data source index keyword in the first sub-retrieval statement is data source: MySQL, database TEST, table test1 (indicating that it can locate the test1 table in the database TEST of MySQL). Query condition: name = test_app. (Indicating to extract records from tset1 that meet the condition that the name field is equal to test_app). Query fields: timefield, uid, ip, id_test (indicating to only extract the fields timefield, uid, ip, and id_test). Deduplication field: host_name (indicating that in the result set, ensure that each value of host_name only appears once, that is, remove duplicate host_name records), limit command: limit the number of retrieval results to 50,000. Rename field: rename host_name to names.
[0057] The data source index keyword in the second sub-retrieval statement is data source: MongoDB, database mongo_test, table log_http (indicating that it can locate the log_http table in the database mongo_test of MongoDB). Query condition: type = ip_address (indicating to extract records from log_http that meet the condition that the type field is equal to ip_address). Statistical operation: count by the field names (count the number of occurrences of each different value of names). In addition, it should be noted that the code logic in the retrieval statement can also represent the execution order of retrievals for each data source. For example, in the above example of the retrieval statement, the data source of the first sub-retrieval statement (i.e., the retrieval statement that appears first in the order of the code) is MySQL, and the data source of the second sub-retrieval statement (i.e., the retrieval statement that appears second in the order of the code) is MongoDB. This indicates that the MySQL data source performs the retrieval first. Here, the retrieval performed by the MySQL data source can be called the first-level query (referring to the initial query). After the retrieval by the MySQL data source is completed, the MongoDB data source performs the retrieval. The retrieval performed by the MongoDB data source can be called the second-level query (i.e., the next query after the initial query). Subsequent query methods can be inferred in this way.
[0058] S103. Convert the retrieval statement into a structured statement according to the keyword.
[0059] The computing device can convert the retrieval statement into a structured statement according to the keywords of the retrieval statement determined in step S102. Converting the retrieval statement into a structured statement is an expression method after classification according to the keywords and the operation objects of the keywords. After converting it into a unified structured statement, the keywords, the operation objects of the keywords, and the conditions after the keywords need to be converted into corresponding structured statements according to the syntax rules of different data sources (a structured statement is a way to express query logic in a standardized and normalized manner. It decomposes complex query statements into a series of clearly defined components (such as keywords, operation objects, and query conditions), and organizes these components in a unified format. This expression method is convenient for understanding and execution, and is also convenient for adapting the query logic to different data sources or query systems). For example, the sub-retrieval statement is "datafilter mysql db=TEST table=test1", the keyword is datafilter, and the operation object follows the keyword. At this time, converting the retrieval statement into a structured statement according to the syntax rules of the mysql data source can be "keyword = `datafilter`, data source = `mysql`, database = `TEST`, data table = `ˋtest1ˋ`", where the keyword datafilter is the data source index operation, mysql is the target data source type, TEST is the database name in the target data source, and test1 is the table name in the TEST database.
[0060] It should be noted that when converting the retrieval statement into a structured statement in the embodiments of the present application, it is also necessary to convert it in combination with conditions such as data source index, retrieval filter conditions, number of retrieval results, and retrieval sorting conditions.
[0061] Exemplarily, for example, the sub-retrieval statement is datafilter mysql db=TEST table=test1·search name=Alic*·filter age>20·order by score DESC·limit1000. Then, converting it into a structured statement can be converted into the following parts:
[0062] Data source index: data source type = `mysql`, database name = `TEST`, table name: = `test1`;
[0063] Retrieval condition: field = `name`, condition = `Alic*(prefix match)`;
[0064] Filter condition: field = `age`, condition = `>20`;
[0065] Sorting condition: field = `score`, sorting method = `DESC (descending order)`.
[0066] S104. Obtain data source parameter information, and convert the structured statement into a retrieval statement corresponding to the data source according to the data source parameter information.
[0067] The computing device can obtain data source parameter information (for example, for MySQL, the data source parameter information can be, for example, the host name, port, user name, and password, etc.). Through the obtained data source parameter information, the data source index in the structured statement can be connected and configured, that is, the placeholder or data source index in the structured statement is extracted. The required parameter values are extracted from the data source parameter information, such as the table name: users, the column name: age, and the value of the query condition: 25, etc. Then, the placeholder or data source index is replaced with the extracted parameter values. The replaced statement is converted into a retrieval statement corresponding to the data source.
[0068] To achieve the purpose of obtaining data from multiple data sources with one retrieval statement. Furthermore, the structured statement can be converted into a retrieval statement corresponding to the data source according to the data source parameter information.
[0069] Since the structured statement disassembles and classifies the retrieval statement according to the retrieval conditions and keywords, the structured statement is converted into the corresponding target retrieval statement according to the data source type of the data source index.
[0070] For example, the sub-retrieval statement "datafilter mysql db=TEST table=test1·search*·limit10" is converted into a structured statement, and the final retrieval statement converted into a Mysql data source can be "SELECT*FROM(SELECT*FROMˋ`test1ˋ`)AS temp_table_1001LIMIT 10".
[0071] In some possible implementation manners, the present application can support the access of multiple data sources, so that users do not need to learn the syntax of each data source. The query of these data sources can be realized with one language provided in the present application. The types of databases that the present application supports as data sources can include domestic databases (infornation technology application innovation adaptation) and Splunk, etc.
[0072] The specific data sources supported can be as follows:
[0073] SQL databases (for example, can include Presto, Hetu, GaussDB, PGSQL, MySQL, Dremio, Hive, and DM, etc.), NoSQL databases (for example, can include MongoDB and ElasticSearch, etc.), in-memory databases (for example, Redis), and streaming data (for example, can include Kafka and MQS).
[0074] Since this application can support access to multiple different types of databases as data sources, and there are significant conceptual differences in the data storage levels of these data sources. To facilitate users to access these heterogeneous data sources in a consistent manner without having to delve into the specific implementation details of each database. This application can uniformly map the storage levels of the data sources into two levels: "database (db)" and "table", that is, store the mapping relationships between the data sources and the db and table of the system. In this scenario, the role of the mapping relationship is to convert the respective unique data storage concepts and structures in different heterogeneous data sources into a unified model that is easy for users to understand and operate (i.e., "database (db)" and "table"). Users can configure the "db" and "table" mapping relationships of the data sources and the connection information of the data sources (such as IP address, port, account, and password, etc.) through the management interface.
[0075] Exemplarily, as shown in Table 1, several mapping relationships between data sources and the db and table of the system are exemplified in Table 1:
[0076] Table 1
[0077]
[0078] The embodiment of this application implements a unified data source access framework, which can support seamless integration of multiple heterogeneous data sources. This not only simplifies the data management process, but also ensures data consistency and integrity, and solves the problem that multiple data sources can perform data management and analysis on a single platform.
[0079] In addition, the retrieval statements provided in the embodiment of this application are more concise, efficient, and have less code compared to the prior art, enabling users to express retrieval requirements in a more intuitive and direct manner, avoiding cumbersome and complex syntax rules and lengthy code. By providing a rich set of built-in functions and command sets, efficient analysis and retrieval of heterogeneous data sources can be achieved.
[0080] Exemplarily, for example, the retrieved user input retrieval statement is:
[0081] datafilter presto db=hive table=test
[0082] Search domain_name=xxx.com
[0083] limit 100
[0084] deduplication client_ip
[0085] fields client_ip, domain_name
[0086] join type = feedback client_ip[search*·stats count by client_ip]
[0087] This retrieval statement represents querying records with domain_name = xxx.com from the test table in the db data source. The returned fields are client_ip and domain_name. The number of results is limited to 100. Deduplication is based on client_ip. Using client_ip as the key value, the results of the initial query are joined with the results of another query. Additionally, statistical operations are performed on client_ip to calculate the occurrence count of each client_ip.
[0088] Among them, datafilter presto db = hive table = test specifies querying specific records from the test table in the db data source, and Search domain_name = xxx.com specifies the filtering condition. Deduplication client_ip clearly specifies deduplication based on client_ip. field client_ip, domain_name directly specifies the returned fields. join type = feedback client_ip[search*·stats count by client_ip] joins the results of the initial query with the results of another query and performs statistics on client_ip. This enables users to quickly build complex query logic using concise code without having to write complex SQL statements or scripts, reducing the time for writing and debugging code, thus enabling efficient analysis and retrieval of heterogeneous data sources.
[0089] S105. Send the retrieval statement to the corresponding data source for retrieval.
[0090] The computing device can send the retrieval statement of the data source obtained through step S104 to the corresponding data source for retrieval. Thus, the final retrieval result can be determined according to the retrieval results of each data source. Furthermore, the cross-data-source joint query (a data operation method that allows users to extract data from multiple data sources and combine these data together according to a certain logic) and in-place computing (a technology that optimizes data processing efficiency, which means that the data processing logic (such as calculation and analysis) is executed as close as possible to the location where the data is stored) are realized. In this application, there is no need to perform data migration, and the data can be directly associated and queried and in-place computed on the original data source (similar to the principle of the association query. For example, after the retrieval statement (also called the query statement) is transformed through steps S102 - S104, it is directly sent to the corresponding data source, so that the data source performs in-place computing according to the retrieval statement), making the data processing more flexible and efficient, and thus improving the data processing efficiency.
[0091] Exemplarily, this application can use a multi-level waterfall join or a multi-level feedback join for cross-data-source joint query. In the multi-level waterfall join, the query passes conditions step by step, and the result of each query will be used as the filtering condition for the next query. The final result includes the result of the last query, that is, execute the first-level query (execute the first query to obtain the result set) and then pass the condition (extract the value of the specified field from the result set of the first-level query as the filtering condition for the second-level query). Among them, the specified field can also be called the target field, which is the field used to locate, filter, or associate data in the data source. For example, the value of a certain ID field can be extracted from the result of the first-level query as the target field. The filtering condition means using the value of the target field as the screening condition to limit the data range returned by the query, and then perform subsequent queries (repeat the above process, passing the result of each level of query to the next level of query until the last query) to obtain the final result (the result set of the last query is used as the final result).
[0092] In a multi-level feedback join, the result of each query is concatenated with the result of the previous query to gradually build the final result set. Each query is executed based on the result of the previous query, and the results are merged. That is, the first-level query is executed (the first query is executed to obtain the result set), then the conditions are passed (the values of the specified fields are extracted from the result set of the first-level query as the filtering conditions for the second-level query), and then the concatenation is performed (the result of the second-level query is concatenated with the result of the first-level query through the specified key value (which can also be called the target key value, such as the primary key or foreign key, such as user_id)), and subsequent queries are performed (the above process is repeated, and the result of each level of query is concatenated with the result of the previous level of query until the last query), and the final result is obtained: (the results of all the gradual concatenations are finally concatenated, and the obtained result is used as the final result).
[0093] This application can access multiple target data sources, then obtain the retrieval statement, split the retrieval statement, determine the keywords of the retrieval statement, and then convert the retrieval statement into the target retrieval statements corresponding to each target data source according to the keywords and the data source parameter information corresponding to each target data source. Furthermore, the target retrieval statements can be sent to the corresponding target data sources respectively, so that the target data sources can use the corresponding target retrieval statements for retrieval. Finally, the target retrieval result can be determined according to the retrieval results of each target data source. Since this application accesses multiple target data sources and can convert the retrieval statement into the target retrieval statements corresponding to each target data source respectively, and then send them to the corresponding target data sources for retrieval. Thus, this method allows users to retrieve data in multiple data sources simultaneously through a unified retrieval statement, without manually switching different systems for retrieval, thereby achieving cross-data-source retrieval and improving the efficiency of data retrieval.
[0094] The following combines specific examples to introduce the data retrieval method of this application, which can be called Embodiment 2 here. In Embodiment 2, retrieving data in two data sources is used as an example, specifically as Figure 2 shown. The data retrieval method includes:
[0095] S201. The processing module of the computing device obtains the input retrieval statement.
[0096] S202. Split the retrieval statement to determine the keywords of the retrieval statement.
[0097] S203. Convert the retrieval statement into a structured statement according to the keywords.
[0098] S204. Convert the structured statement into the retrieval statement corresponding to Data Source A.
[0099] Convert the structured statement into a retrieval statement corresponding to data source A. For example, data source A can be a mysql database here.
[0100] S205. Convert the structured statement into a retrieval statement corresponding to data source B.
[0101] Convert the structured statement into a retrieval statement corresponding to data source B. For example, data source B can be a mongo database here.
[0102] The processing module can construct an execution queue according to the retrieval statement input by the user and the converted retrieval statement (which can also be called the target retrieval statement here), determine the execution order of the target retrieval statement, so that the data source (which can also be called the target data source here) executes the corresponding retrieval statement according to this execution order, and then execute step S206 and step S207.
[0103] S206. Send the retrieval statement corresponding to data source A to data source A for retrieval.
[0104] Send the retrieval statement corresponding to data source A in the execution queue to data source A for retrieval.
[0105] S207. Send the retrieval statement corresponding to data source B to data source B for retrieval.
[0106] Send the retrieval statement corresponding to data source B in the execution queue to data source B for retrieval.
[0107] The implementation principles of steps S201 - S207 are similar to those of steps S101 - S104 in the first embodiment. For specific details, please refer to the corresponding content in the first embodiment, and no redundant description will be given here.
[0108] S208. Determine the final retrieval result according to the retrieval results of each data source.
[0109] The processing module can determine the final retrieval result according to the retrieval results of each data source.
[0110] Among them, the left query (that is, the retrieval by data source A in step S206, which can also be called a left join query) can return the hosts with records in the past preset time period (for example, half an hour), and the right query (that is, the retrieval by data source B in step S207, which can also be called a right join query) returns the list of machines with ip records. In this application, through the "feedback mode", batch joins can be automatically performed inside the association type, and finally aggregated into a complete set, thereby supporting the association of large tables with millions of rows.
[0111] It should be noted that the essence of the join command is the corresponding left and right subqueries (here taking the retrieval in data source A and data source B as an example). The head command can be used to specify the maximum amount of data participating in the join respectively. If the left and right queries of the join are both queries of hetu / presto homogeneous data sources, and all the commands therein can be pushed down to hetu / presto for execution, then the left and right queries and the join can all be pushed down to hetu (a SQL query engine) for execution at one time. In this case, the head command can be used to specify that up to 10 million data each participate in the join for the left and right subqueries respectively.
[0112] In some possible implementation manners, when the cross-data-source join type is waterfall, the values of the specified fields in the left query result set can be passed to the right query as the query filtering conditions of the right query, and the result set returned by the right query is used as the final result set.
[0113] When the join type is feedback, the values of the specified fields in the left query result set can be passed to the right query as the query filtering conditions of the right query, and the result set returned by the right query is then associated and spliced with the records with the same key values in the left query. The spliced result set is used as the final result set. In this way, the final retrieval result can be obtained.
[0114] The following introduces the hardware structure of the computing device of the present application. Refer to Figure 3 , Figure 3 which is a schematic diagram of the hardware structure of a computing device provided by an embodiment of the present application.
[0115] As Figure 3 shown, the computing device 1000 includes a processor 1010 and a memory 1020; wherein, the memory 1020 stores computer instructions, and the processor 1010 is used to execute the computer instructions, so that the computing device 1000 executes the data retrieval method shown above.
[0116] In some embodiments, the processor 1010 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0117] In some embodiments, the memory 1020 may be a volatile memory or a non-volatile memory, such as registers. Specifically, a volatile memory refers to a memory in which the data stored internally is lost when the power supply is interrupted. Among them, the volatile memory is mainly a random access memory (RAM), including a static random access memory (SRAM) and a dynamic random access memory (DRAM). A non-volatile memory refers to a memory in which the data stored internally is not lost even when the power supply is interrupted. Common non-volatile memories include read only memory (ROM), optical discs, magnetic disks, solid state drives, and various memory cards based on flash memory technology.
[0118] In some embodiments, the memory 1020 has executable code, and the memory 1010 executes this code to implement the search for the server model.
[0119] The communication interface 1030 is used for external communication. For example, the communication interface 1030 serves as a first interface to implement communication with real devices.
[0120] The bus may be a Peripheral Component Interconnect (PCI) bus, an extended industry standard architecture (eisa) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of understanding, Figure 3 it is represented by only one thick line, but it does not mean that there is only one bus or one type of bus.
[0121] Embodiments of the present application also provide a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on a computing device, it causes the computing device to execute the above data retrieval method. Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above data retrieval method.
[0122] In some possible implementation manners, the computing device may be a server or a terminal, etc. Of course, this is only an example here and is not subject to any limitation.
[0123] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.
[0124] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data retrieval method, characterized in that, Applied to a computing device that accesses multiple target data sources, the method includes: Obtain a retrieval statement; Split the retrieval statement to determine the keywords of the retrieval statement; Generate target retrieval statements respectively corresponding to each target data source according to the keywords and the data source parameter information respectively corresponding to each target data source; Send the target retrieval statements to the corresponding target data sources respectively, so that the target data sources use the corresponding target retrieval statements for retrieval; Determine the target retrieval result according to the retrieval results of each target data source.
2. The method according to claim 1, wherein The determining the target retrieval result according to the retrieval results of each target data source includes: Extract the values of the target fields from the retrieval results of each level of target data sources as the filtering conditions for querying the next level of target data sources until the retrieval results of the last level of target data sources are obtained, where the target fields are fields used to locate, filter or associate data in the data source; Determine the retrieval result of the last level of target data source as the target retrieval result.
3. The method according to claim 1, characterized in that The determining the target retrieval result according to the retrieval results of each target data source includes: Extract the values of the target fields from the retrieval results of each level of target data sources as the filtering conditions for querying the next level of target data sources; Concatenate the retrieval results of the next level of target data sources with the retrieval results of the previous level of target data sources through the target key values until the retrieval results of the last level of target data sources are obtained; Concatenate the results obtained by the associated concatenation to determine the target retrieval result.
4. The method according to any one of claims 1 to 3, characterized in that, The splitting the retrieval statement to determine the keywords of the retrieval statement includes: Split the retrieval statement according to the data source index keyword and the target data syntax command keyword to obtain the keywords of the sub-retrieval statements corresponding to each data source.
5. The method according to claim 4, wherein The generating the target retrieval statements respectively corresponding to each target data source includes: According to the keywords of the sub-retrieval statements, the operation objects of the keywords, and the conditions after the keywords, convert each sub-retrieval statement into a structured statement according to the syntax rules corresponding to the data source; Convert the structured statement into the target retrieval statement corresponding to the target data source.
6. The method according to claim 5, wherein The structured statement includes at least one of a data source index, a retrieval condition, a filtering condition, a statistical function, an aggregation function, a non-aggregation function, a merge command, a sorting command, a command to retrieve the first N data records, and a multi-data-source fusion command. The data source index is used to represent the source of data in a sub-retrieval statement. The retrieval condition is used to filter out records that meet the conditions in the data source of the sub-retrieval statement. The filtering condition is used to remove duplicate records and filter specific fields in the sub-retrieval statement. The statistical function is used to perform statistical analysis on the data in the sub-retrieval statement. The aggregation function is used to merge multiple rows of data in the sub-retrieval statement into a single row of data. The non-aggregation function is used to operate on a single row of data in the sub-retrieval statement. The merge command is used to merge multiple query results into a single result set. The sorting command is used to sort the query results. The command to retrieve the first N data records is used to limit the number of query results. The multi-data-source fusion command is used to retrieve data from multiple data sources and merge the results into a single result set.
7. The method according to claim 5, characterized in that, Converting the structured statement into a target retrieval statement corresponding to a target data source includes: extracting placeholder or data source index in the structured statement; extracting the required parameter values from the corresponding data source parameter information, and replacing the placeholder or data source index with the extracted parameter values; converting the statement after replacement into a target retrieval statement corresponding to the data source.
8. The method according to claim 1, wherein The method further includes: mapping the storage levels of different target data sources to databases and tables to obtain the mapping relationship corresponding to each data source, where the mapping relationship is used to map the respective corresponding storage levels in different heterogeneous data sources to a unified level; storing the mapping relationship.
9. The method according to claim 1, characterized in that, Sending the target retrieval statement to the corresponding target data source respectively includes: constructing an execution queue according to the target retrieval statement, where the execution queue is used to determine the execution order of each target retrieval statement; sending the target retrieval statement to the corresponding target data source respectively from the execution queue.
10. A computing device, characterized in that, including a memory and a processor; the memory is coupled with the processor; the memory stores program instructions, and when the program instructions are executed by the processor, the computing device executes the method according to any one of claims 1-9.