Data query method and apparatus

By extracting query conditions in the HBase database and storing them in the target data lake, the problem of poor performance of HBase database when processing non-primary key field queries is solved, and the effect of improving query efficiency is achieved.

WO2025118788A1PCT designated stage expired Publication Date: 2025-06-12CHINA TELECOM CORP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/120746
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-09-24
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

When HBase database processes queries with non-primary key fields, the lack of indexes leads to full table scanning, resulting in poor query performance.

Method used

By obtaining data query requests, query conditions are extracted, and data is extracted from the target database when the query conditions belong to the preset primary key. When the query conditions belong to the non-preset primary key, data is extracted from the target data lake, and the storage performance advantages of the data lake are used to improve query efficiency.

Benefits of technology

By storing the data corresponding to non-preset primary keys in the target database in the target data lake, and leveraging the performance advantages of the data lake, the technical effect of improving the data query efficiency of the target database is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120746_12062025_PF_FP_ABST
    Figure CN2024120746_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a data query method and apparatus. The method comprises: acquiring a data query request, and extracting a query condition from the data query request; when the query condition belongs to a preset primary key, extracting from a target database data corresponding to the query condition, and when the query condition belongs to a non-preset primary key, extracting from a target data lake data corresponding to the query condition, wherein the target database is a database which stores data in the form of key-value pairs, and the target data lake stores data in the target database that corresponds to the query condition not belonging to the preset primary key; and outputting the data corresponding to the query condition.
Need to check novelty before this filing date? Find Prior Art

Description

Data query method and device

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 202311667498.2, filed on December 6, 2023, entitled “Data Query Method and Device,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of database technology, and in particular to a data query method and device. Background Art

[0004] With the rapid development of big data, the diversity and complexity of business data have gradually increased. A single relational database can no longer meet all business needs, and data lakes have emerged. A data lake is a centralized storage database that allows the storage of all structured and unstructured data at any scale and can run various types of queries and analyses based on relational databases. HBase is a database that uses Hadoop HDFS as its storage base for massive data storage, storing data in key-value pairs, using the row key as the primary key. HBase supports query and analysis based on the query engine, but due to the storage characteristics of HBase, when processing queries on non-primary key fields, these fields lack indexes, often requiring an HBase full table scan, resulting in poor query performance.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a data query method and device to at least solve the technical problem of poor key-value database query performance in related technologies.

[0007] In a first aspect, the present application provides a data query method, including: obtaining a data query request and extracting a query condition from the data query request; when the query condition belongs to a preset primary key, extracting data corresponding to the query condition from a target database; when the query condition belongs to a non-preset primary key, extracting data corresponding to the query condition from a target data lake, where the target database is a database that stores data in the form of key-value pairs, and the target data lake stores data in the target database corresponding to the query condition that does not belong to the preset primary key; and outputting the data corresponding to the query condition.

[0008] In some embodiments, determining whether the query condition belongs to the preset primary key includes: obtaining at least one subkey contained in the query condition; combining at least one subkey to obtain a target key; and determining that the query condition belongs to the preset primary key when the target key is the same as the preset primary key or the target key is a prefix of the preset primary key.

[0009] In some embodiments, extracting data corresponding to the query condition from the target database includes: extracting a target value corresponding to the query condition when the target key is the same as the preset primary key; determining a row key based on the target value, and obtaining the data corresponding to the query condition through an acquisition interface of the target database according to the row key.

[0010] In some embodiments, the above method also includes: extracting a target value corresponding to the query condition when the target key is a prefix of a preset primary key; determining a query range of the query condition based on the target value, and obtaining data corresponding to the query condition according to the query range.

[0011] In some embodiments, before extracting data corresponding to the query conditions from the target data lake, the above method also includes: performing snapshot processing on the data table to be synchronized in the target database to obtain snapshot data, and recording snapshot information of the data table to be synchronized, where the snapshot data is the full data of the data table to be synchronized; extracting the current maximum sequence number of each data unit in the data table to be synchronized from the snapshot information; and writing the snapshot data into the target data lake until the sequence number of the written data is equal to the current maximum sequence number, determining that the snapshot data writing is complete.

[0012] In some embodiments, after the snapshot data is written, the above method includes: obtaining incremental data and incremental data logs, and sorting the incremental data logs; converting the incremental data logs into a data table format in the target data lake, and writing the incremental data into the target data lake through the message middleware in the order of the incremental data logs; obtaining the data unit sequence number in the incremental data written into the target data lake, and determining the writing progress of the incremental data based on the data unit sequence number; and determining the writing delay of the incremental data based on the writing progress.

[0013] In some embodiments, when the query condition belongs to a non-preset primary key, extracting data corresponding to the query condition from the target data lake includes: obtaining the data sending progress of the message middleware; determining the sending progress and the writing progress as the write delay of the incremental data; and extracting the data corresponding to the query condition from the target data lake when the write delay is less than a preset threshold.

[0014] In a second aspect, the present application further provides a data query device, including: an acquisition module, used to obtain a data query request and extract query conditions from the data query request; an extraction module, used to extract data corresponding to the query conditions from a target database when the query conditions belong to a preset primary key; when the query conditions belong to a non-preset primary key, extracting data corresponding to the query conditions from a target data lake, where the target database is a database that stores data in the form of key-value pairs, and the target data lake stores data corresponding to the query conditions in the target database that belong to the non-preset primary key; and an output module, used to output the data corresponding to the query conditions.

[0015] In some embodiments, the extraction module also includes a verification submodule, which is used to: obtain at least one subkey contained in the query condition; combine the at least one subkey to obtain a target key; and determine that the query condition belongs to the preset primary key when the target key is the same as the preset primary key or the target key is a prefix of the preset primary key.

[0016] In some embodiments, the verification submodule includes a first extraction unit and a second extraction unit; the first extraction unit is used to: extract the target value corresponding to the query condition when the target key is the same as the preset primary key; determine the row key based on the target value, and obtain the data corresponding to the query condition through the acquisition interface of the target database according to the row key; the second extraction unit is used to: extract the target value corresponding to the query condition when the target key is a prefix of the preset primary key; determine the query range of the query condition based on the target value, and obtain the data corresponding to the query condition according to the query range.

[0017] In some embodiments, the extraction module also includes a synchronization submodule, which is used to: perform snapshot processing on the data table to be synchronized in the target database to obtain snapshot data, and record the snapshot information of the data table to be synchronized, where the snapshot data is the full data of the data table to be synchronized; extract the current maximum sequence number of each data unit in the data table to be synchronized from the snapshot information; write the snapshot data into the target data lake until the sequence number of the written data is equal to the current maximum sequence number, and determine that the snapshot data writing is completed.

[0018] In some embodiments, the synchronization submodule further includes a synchronization unit, which is used to: obtain incremental data and incremental data logs, and sort the incremental data logs; convert the incremental data logs into the data table format in the target data lake, and write the incremental data into the target data lake through the message middleware in the order of the incremental data logs; obtain the data unit sequence number in the incremental data written into the target data lake, and determine the writing progress of the incremental data based on the data unit sequence number; determine the writing delay of the incremental data based on the writing progress.

[0019] In some embodiments, the synchronization unit further includes a determination subunit, which is used to: obtain the data sending progress of the message middleware; determine the sending progress and the writing progress as the write delay of the incremental data; and extract the data corresponding to the query condition from the target data lake when the write delay is less than a preset threshold.

[0020] In a third aspect, the present application further provides a non-volatile computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the data query method in any embodiment of the first aspect is executed.

[0021] In a fourth aspect, the present application further provides a computer device comprising: a memory and a processor. The memory stores a computer program, and the processor executes the data query method in any embodiment of the first aspect when running the computer program.

[0022] According to the data query method and device provided in the embodiment of the present application, a data query request is obtained and query conditions are extracted from the data query request; when the query conditions belong to a preset primary key, the data corresponding to the query conditions are extracted from the target database; when the query conditions belong to a non-preset primary key, the data corresponding to the query conditions are extracted from the target data lake, the target database is a database that stores data in the form of key-value pairs, and the target data lake stores data in the target database that does not correspond to the query conditions of the preset primary key; and the data corresponding to the query conditions is output. Thus, by storing the data corresponding to the query conditions of the target database that do not correspond to the preset primary key in the target data lake, and when the query conditions do not belong to the preset primary key, the storage performance advantage of the data lake is used to query the corresponding data, the dual advantages of the target database and the target data lake are fully utilized, achieving the technical effect of improving the efficiency of target database data query, and thus solving the technical problem of poor key-value database query performance in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be derived from these drawings without inventive effort.

[0024] FIG1 is a hardware structure block diagram of a computer device (or mobile terminal) for a data query method according to an embodiment of the present application.

[0025] FIG2 is a flow chart of a data query method according to an embodiment of the present application.

[0026] FIG3 is a flow chart of a data query method according to another embodiment of the present application.

[0027] FIG4 is a flow chart of a data query method according to another embodiment of the present application.

[0028] FIG5 is a schematic structural diagram of a data query device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer device or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer device (or mobile terminal) for implementing a data query method. As shown in Figure 1, the computer device 10 (or mobile terminal) may include one or more (as shown in 102a, 102b, ..., 102n in Figure 1) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor (MCU) or a programmable logic device (FPGA)), a memory 104 for storing data, and a transmission module 106 for communication. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be used as a port of a bus (BUS)), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than those shown in Figure 1, or have a configuration different from that shown in Figure 1.

[0033] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer device 10 (or mobile terminal). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0034] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data query method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizes the above-mentioned data query method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the computer device 10 via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0035] The transmission device 106 is used to receive or send data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer device 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0036] The display may be, for example, a touch screen liquid crystal display (LCD), which may enable a user to interact with a user interface of the computer device 10 (or mobile terminal).

[0037] In the above operating environment, an embodiment of the present application provides a data query method, as shown in FIG2 , which includes the following steps S202 to S206 .

[0038] Step S202: Obtain a data query request and extract query conditions from the data query request.

[0039] In step S204, when the query condition belongs to a preset primary key, data corresponding to the query condition is extracted from the target database; when the query condition belongs to a non-preset primary key, data corresponding to the query condition is extracted from the target data lake; the target database is a database that stores data in the form of key-value pairs, and the target data lake stores data in the target database corresponding to query conditions that do not belong to the preset primary key.

[0040] Step S206: output data corresponding to the query condition.

[0041] Through the above steps, it is possible to obtain a data query request and extract the query conditions from the data query request; when the query conditions belong to the preset primary key, extract the data corresponding to the query conditions from the target database; when the query conditions belong to non-preset primary keys, extract the data corresponding to the query conditions from the target data lake. The target database is a database that stores data in the form of key-value pairs, and the target data lake stores data in the target database that does not correspond to the preset primary key for the query conditions; and output the data corresponding to the query conditions. In this embodiment, by storing the data in the target database that corresponds to the non-preset primary key for the query conditions in the target database in the target data lake, and when the query conditions do not belong to the preset primary key, the storage performance advantage of the data lake is used to query the corresponding data, thereby fully utilizing the dual advantages of the target database and the target data lake, achieving the technical effect of improving the efficiency of target database data queries, and thereby solving the technical problem of poor key-value database query performance in related technologies.

[0042] It should be noted that the target data lake can adopt a data lake architecture (for example, Hudi). Before the query begins, the data in the target database can be synchronized to the target data lake. When performing data queries, a query engine (for example, Trino) can be used for the query, and the target database can be an HBase database.

[0043] In actual application scenarios, when executing data queries based on Trino, Trino will push down the query conditions and determine whether it is a primary key query based on the query conditions. If it is a primary key query, the data will be queried directly through HBase; if it is a non-primary key query, it will determine whether the delay threshold configuration is met. If so, the data will be queried from Hudi, otherwise, the data will be queried based on HBase.

[0044] Steps S202 to S206 are described below through a specific embodiment.

[0045] In some embodiments of the present application, the step of determining whether the query condition belongs to the preset primary key may include: obtaining at least one subkey contained in the query condition; combining at least one subkey to obtain a target key; and determining that the query condition belongs to the preset primary key when the target key is the same as the preset primary key or the target key is a prefix of the preset primary key.

[0046] Taking the preset primary key (custom_id, commodity_id) as an example, the subkeys included in the query condition are determined. If the order of the subkeys included in the query condition is custom_id and commodity_id, the target key is considered to be "(custom_id, commodity_id)", and it can be determined that the target key is the same as the preset primary key; if the query condition only contains the subkey custom_id, the target key is determined to be custom_id, which is a prefix of the preset primary key, and it can be determined that the query condition belongs to the preset primary key; if the subkey included in the query condition is "commodity_id", the target key is determined to be "commodity_id", which is not a prefix of the preset primary key, and the query condition does not belong to the preset primary key.

[0047] It is understandable that if the subkeys in the query condition include the first key in the preset primary key and the order of the subkeys in the target key is the same as the order of the preset primary key, the target key is considered to be a prefix of the preset primary key.

[0048] There are two different ways to extract data corresponding to the query conditions from the target database:

[0049] Method 1 is: when the target key is the same as the preset primary key, extract the target value corresponding to the query condition; determine the row key based on the target value, and obtain the data corresponding to the query condition through the acquisition interface of the target database based on the row key.

[0050] Method 2: When the target key is a prefix of the preset primary key, extract the target value corresponding to the query condition; determine the query scope based on the target value, and retrieve the data corresponding to the query condition based on the query scope. In a real-world scenario, for example, to query order details for products with commodity_ids 212 and 324, customers with custom_ids 102 and 203 would determine the target values ​​corresponding to the query condition as 102, 203, 212, and 324; row keys can be generated based on the Cartesian product, resulting in four row keys: 102_212, 102_324, 203_212, and 203_324. Then, use the HBase Get interface to retrieve the order data corresponding to the row keys.

[0051] Taking the example of obtaining orders for customer with custom_id 102, we can directly generate the query interval [102,103), and then scan the row keys of the query interval in the HBase database to obtain all order data corresponding to customer 102.

[0052] For example, to obtain the sales status of the product with commodity_id 212, since the query condition starts with commodity_id, we cannot use the order of HBase's RowKey, nor can we query directly through the Get interface. A full table scan of the data in the HBase database is required, and the query request must be routed to the target data lake to obtain the desired data.

[0053] Before extracting data corresponding to the query conditions from the target data lake, the above method also includes: performing snapshot processing on the data table to be synchronized in the target database to obtain snapshot data, and recording snapshot information of the data table to be synchronized, where the snapshot data is the full data of the data table to be synchronized; extracting the current maximum sequence number of each data unit in the data table to be synchronized from the snapshot information; and writing the snapshot data to the target data lake until the sequence number of the written data is equal to the current maximum sequence number, determining that the snapshot data writing is complete.

[0054] It should be noted that when the sequence number of the written data of each data unit is equal to the maximum sequence number, it is determined that the snapshot data writing is completed.

[0055] Specifically, step 1: obtain the structure and definition information of objects such as tables, columns, keys, indexes, etc. in the HBase database from Trino, create a table with the same name in the target data lake as the data lake result table, and set the preset threshold corresponding to the write delay. Step 2: Use a custom ReplicationEndpoint (a terminal node for transmitting and receiving replicated data) for the data table to be synchronized to obtain the write-ahead log (WAL) of the incremental data. By writing a log plug-in, these WAL logs are sent to the message middleware for storage based on the ReplicationEndpoint, and the log plug-in, the data table to be synchronized, and the column cluster information are configured at the same time; the log plug-in will report the current data update progress.<region,sequence_id> The metadata management module records the progress of the data transmission. Step 3: Full synchronization: Take a snapshot of the data table to be synchronized and record the snapshot information. From the snapshot, extract the current maximum sequence_id (sequence number) in each region (data unit) and report the snapshot information to the metadata management module. Concurrently, write the snapshot data to the target data lake result table.

[0056] After the snapshot data is written, the method further includes: obtaining incremental data and incremental data logs, and sorting the incremental data logs; converting the incremental data logs into a data table format in the target data lake, and writing the incremental data into the target data lake through the message middleware in the order of the incremental data logs; obtaining the data unit sequence number in the incremental data written into the target data lake, and determining the writing progress of the incremental data according to the data unit sequence number; and determining the writing delay of the incremental data according to the writing progress.

[0057] Specifically, after the full synchronization task is completed, the incremental log data is pulled from the log management module, the data is consumed from the message middleware and the data is filtered based on the synchronization progress in the full synchronization, the WAL log is sorted by window, parsed and converted into the data lake table format and written to the target data lake, and the data written at this time is recorded.<region,sequence_id> The write progress is recorded as the write delay, and the difference between the write progress and the send progress is recorded as the write delay.

[0058] When the query condition belongs to a non-preset primary key, the step of extracting data corresponding to the query condition from the target data lake includes: obtaining the data sending progress of the message middleware; determining the sending progress and the writing progress as the write delay of the incremental data; and when the write delay is less than a preset threshold, extracting the data corresponding to the query condition from the target data lake.

[0059] Taking the preset primary key "(custom_id, commodity_id)" as an example, Figure 3 shows an optional data query method. As shown in Figure 3, first determine whether the query conditions contain query conditions with the same value as the preset primary key custom_id (i.e., the same query conditions), and then determine whether the query conditions contain queries with the same value as the preset primary key commodity_id. If both conditions are met, a Cartesian connection is generated based on the value corresponding to the preset primary key, and a row key is generated. Based on the row key, the corresponding data is queried in the HBase database. If custom_id is not included, determine whether the synchronization delay (i.e., write delay) is less than a preset threshold (e.g., a configuration threshold). If it is less than the preset threshold, query the data in the Hudi data lake.

[0060] Figure 4 shows another data query method. As shown in Figure 4, when a query request is received, data is queried in the HBase database using the primary key query method, and data is queried in the Hudi data lake using other query methods. When an update request is received, the data is updated. The HBase database performs incremental data synchronization with the Hudi data lake.

[0061] The present application embodiment provides a data query device, as shown in FIG5 , including:

[0062] The acquisition module 50 is configured to acquire a data query request and extract query conditions from the data query request.

[0063] Extraction module 52 is used to: extract data corresponding to the query condition from the target database when the query condition belongs to a preset primary key; extract data corresponding to the query condition from the target data lake when the query condition belongs to a non-preset primary key; the target database is a database that stores data in the form of key-value pairs, and the target data lake stores data corresponding to the query condition in the target database that belongs to a non-preset primary key.

[0064] The output module 54 is used to output data corresponding to the query conditions.

[0065] The extraction module 52 also includes a verification submodule, which is used to: obtain at least one subkey contained in the query condition; combine at least one subkey to obtain a target key; and determine that the query condition belongs to the preset primary key when the target key is the same as the preset primary key or the target key is a prefix of the preset primary key.

[0066] In some embodiments, the verification submodule may further include: a first extraction unit and a second extraction unit. The first extraction unit is configured to extract a target value corresponding to the query condition when the target key is the same as the preset primary key; determine a row key based on the target value, and obtain data corresponding to the query condition through an acquisition interface of the target database based on the row key. The second extraction unit is configured to extract the target value corresponding to the query condition when the target key is a prefix of the preset primary key; determine a query range of the query condition based on the target value, and obtain data corresponding to the query condition based on the query range.

[0067] The extraction module 52 also includes a synchronization submodule, which is used to: perform snapshot processing on the data table to be synchronized in the target database to obtain snapshot data, and record the snapshot information of the data table to be synchronized, where the snapshot data is the full data of the data table to be synchronized; extract the current maximum sequence number of each data unit in the data table to be synchronized from the snapshot information; write the snapshot data into the target data lake until the sequence number of the written data is equal to the current maximum sequence number, and determine that the snapshot data writing is completed.

[0068] In some embodiments, the synchronization submodule may also include a synchronization unit, which is used to: obtain incremental data and incremental data logs, and sort the incremental data logs; convert the incremental data logs into the data table format in the target data lake, and write the incremental data into the target data lake through the message middleware in the order of the incremental data logs; obtain the data unit sequence number in the incremental data written into the target data lake, and determine the writing progress of the incremental data based on the data unit sequence number; determine the writing delay of the incremental data based on the writing progress.

[0069] In some embodiments, the synchronization unit may also include a determination subunit for: obtaining the data sending progress of the message middleware; determining the sending progress and the writing progress as the write delay of the incremental data; and extracting the data corresponding to the query condition from the target data lake when the write delay is less than a preset threshold.

[0070] According to another aspect of the present application, a non-volatile computer-readable storage medium is provided, having a computer program stored thereon. When executed by a processor, the computer program executes the data query method, including: obtaining a data query request and extracting a query condition from the data query request; if the query condition belongs to a preset primary key, extracting data corresponding to the query condition from a target database; if the query condition belongs to a non-preset primary key, extracting data corresponding to the query condition from a target data lake, wherein the target database stores data in the form of key-value pairs, and the target data lake stores data in the target database corresponding to the query condition that does not belong to the preset primary key; and outputting the data corresponding to the query condition.

[0071] According to another aspect of the present application, a computer device is provided, comprising: a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method for performing the data query described above is performed, including: obtaining a data query request and extracting a query condition from the data query request; if the query condition corresponds to a preset primary key, extracting data corresponding to the query condition from a target database; if the query condition corresponds to a non-preset primary key, extracting data corresponding to the query condition from a target data lake, wherein the target database stores data in the form of key-value pairs, and the target data lake stores data in the target database corresponding to the query condition that does not correspond to the preset primary key; and outputting the data corresponding to the query condition.

[0072] It should be noted that the various modules in the above-mentioned data query device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0073] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0074] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0075] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0076] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.

[0077] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0078] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0079] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0080] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A data query method, comprising: Obtaining a data query request, and extracting a query condition from the data query request; When the query condition belongs to a preset primary key, extracting data corresponding to the query condition from the target database; In the case where the query condition belongs to a non-preset primary key, data corresponding to the query condition is extracted from the target data lake, the target database is a database for storing data in key-value pairs, and the target data lake stores data in the target database where the query condition does not correspond to the preset primary key; as well as Output the data corresponding to the query condition.

2. The method according to claim 1, wherein determining whether the query condition belongs to a preset primary key comprises: Obtain at least one subkey included in the query condition; Combining the at least one subkey to obtain a target key; In the case that the target key is the same as the preset primary key or the target key is a prefix of the preset primary key, it is determined that the query condition belongs to the preset primary key.

3. The method according to claim 2, wherein extracting data corresponding to the query condition from the target database comprises: In the case where the target key is the same as the preset primary key, extracting the target value corresponding to the query condition; A row key is determined based on the target value, and data corresponding to the query condition is acquired through an acquisition interface of the target database according to the row key.

4. The method according to claim 2, further comprising: In the case where the target key is a prefix of the preset primary key, extracting a target value corresponding to the query condition; A query range of the query condition is determined based on the target value, and data corresponding to the query condition is acquired according to the query range.

5. The method according to claim 1, wherein before extracting the data corresponding to the query condition from the target data lake, the method further comprises: Performing snapshot processing on the data table to be synchronized in the target database to obtain snapshot data, and recording snapshot information of the data table to be synchronized, wherein the snapshot data is the full amount of data of the data table to be synchronized; Extracting the current maximum sequence number of each data unit in the to-be-synchronized data table from the snapshot information; The snapshot data is written into the target data lake until the sequence number of the written data is equal to the current maximum sequence number, and then it is determined that the writing of the snapshot data is completed.

6. The method according to claim 5, wherein after the snapshot data is written, the method comprises: Acquire incremental data and incremental data logs, and sort the incremental data logs; Convert the incremental data log into a data table format in the target data lake, and write the incremental data into the target data lake through the message middleware in the order of the incremental data log; Obtaining a data unit sequence number in the incremental data written into the target data lake, and determining a writing progress of the incremental data according to the data unit sequence number; A write delay for the incremental data is determined according to the write progress.

7. The method according to claim 6, wherein when the query condition belongs to a non-preset primary key, extracting data corresponding to the query condition from the target data lake comprises: Obtaining the data sending progress of the message middleware; Determine the sending progress and the writing progress as a writing delay of the incremental data; When the write delay is less than a preset threshold, data corresponding to the query condition is extracted from the target data lake.

8. A data query device, comprising: An acquisition module, used to acquire a data query request and extract a query condition from the data query request; An extraction module, used for extracting data corresponding to the query condition from the target database when the query condition belongs to a preset primary key; In the case where the query condition belongs to a non-preset primary key, data corresponding to the query condition is extracted from a target data lake, the target database is a database that stores data in the form of key-value pairs, and the target data lake stores data corresponding to the query condition in the target database that belongs to the non-preset primary key; The output module is used to output the data corresponding to the query condition.

9. The device according to claim 8, wherein the extraction module further comprises a verification submodule, wherein the verification submodule is used to: Obtain at least one subkey included in the query condition; Combining the at least one subkey to obtain a target key; In the case that the target key is the same as the preset primary key or the target key is a prefix of the preset primary key, it is determined that the query condition belongs to the preset primary key.

10. The device according to claim 9, wherein the verification submodule comprises a first extraction unit and a second extraction unit; The first extraction unit is used to: extract the target value corresponding to the query condition when the target key is the same as the preset primary key; determine the row key based on the target value, and obtain the data corresponding to the query condition through the acquisition interface of the target database according to the row key; The second extraction unit is used to: extract the target value corresponding to the query condition when the target key is a prefix of the preset primary key; determine the query range of the query condition based on the target value, and obtain data corresponding to the query condition according to the query range.

11. The device according to claim 8, wherein the extraction module further comprises a synchronization submodule, wherein the synchronization submodule is configured to: Performing snapshot processing on the data table to be synchronized in the target database to obtain snapshot data, and recording snapshot information of the data table to be synchronized, wherein the snapshot data is the full amount of data of the data table to be synchronized; Extracting the current maximum sequence number of each data unit in the to-be-synchronized data table from the snapshot information; The snapshot data is written into the target data lake until the sequence number of the written data is equal to the current maximum sequence number, and then it is determined that the writing of the snapshot data is completed.

12. The device according to claim 11, wherein the synchronization submodule further comprises a synchronization unit, wherein the synchronization unit is configured to: Acquire incremental data and incremental data logs, and sort the incremental data logs; Convert the incremental data log into a data table format in the target data lake, and write the incremental data into the target data lake through the message middleware in the order of the incremental data log; Obtaining a data unit sequence number in the incremental data written into the target data lake, and determining a writing progress of the incremental data according to the data unit sequence number; A write delay for the incremental data is determined according to the write progress.

13. The device according to claim 12, wherein the synchronization unit further comprises a determination subunit, wherein the determination subunit is configured to: Obtaining the data sending progress of the message middleware; Determine the sending progress and the writing progress as a writing delay of the incremental data; When the write delay is less than a preset threshold, data corresponding to the query condition is extracted from the target data lake.

14. A non-volatile computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, executes the data query method according to any one of claims 1 to 7.

15. A computer device comprising: A memory and a processor, wherein the memory stores a computer program, and the processor executes the data query method according to any one of claims 1 to 7 when running the computer program.

Citation Information

Patent Citations

  • Data query method, device and equipment based on Hudi and storage medium

    CN113094340A

  • Data lake data processing method and system

    CN115185955A

  • Data query method and device, electronic equipment and storage medium

    CN115374155A

  • Data information query method and device, equipment and storage medium

    CN116431673A

  • Data query method and device

    CN117931936A