Data query method and system, electronic equipment, storage medium and program product

By receiving and parsing query requests in real-time data warehouses and acquiring data in the second storage component using the mapping table, the problems of limited data query functions and low efficiency in the prior art are solved, and more efficient data query is achieved.

CN120123374APending Publication Date: 2025-06-10KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510362035.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the field of real-time data warehouses, the existing technology can only perform data query through Offset, and cannot filter according to the content of the field, resulting in difficulty in positioning problems, inefficient inspection, and difficult to meet users' data query needs.

Method used

By receiving a query request for the service data in the first storage component of the service system, analyzing the query request, querying the preset mapping table, and obtaining the second storage unit corresponding to the first storage unit, thereby obtaining the query result. The second storage component provides different data query functions to meet the user's data query needs.

Benefits of technology

It realizes the data query capability that meets user needs while ensuring timeliness, improves data query efficiency, and solves the problems of limited data query functions, poor timeliness and low data query efficiency in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123374A_ABST
    Figure CN120123374A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data query method and system, electronic equipment, a storage medium and a program product. The method comprises the following steps: receiving a query request for business data in a first storage component of a business system; analyzing the query request to obtain a first storage unit of the business data in the first storage component; querying a preset mapping table to obtain a second storage unit corresponding to the first storage unit; the second storage unit is a storage unit which is correspondingly created in the second storage component and is used for storing service data; obtaining a query result corresponding to the query request from a second storage unit; wherein the first storage component and the second storage component respectively provide different data query functions, so that the second storage unit can be used as a redundant backup unit of service data, the main process of service processing is still based on the first storage component, the data query capability better meeting user requirements can be provided, and the data query efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a data query method, system, electronic device, storage medium, and program product. Background Art

[0002] In the field of real-time data warehouses, business data is collected from business databases to message queues and real-time metrics are generated after being consumed by real-time engines. However, usually, the storage components in business systems may only support a certain specific data query method and can only provide limited data query functions. For example, the message queue data mainly based on Kafka only supports querying through Offset and cannot be filtered according to field content. Therefore, when data quality problems occur, it is difficult to locate the problems and the troubleshooting efficiency is extremely low, making it difficult to meet the data query needs of users. Summary of the Invention

[0003] In view of the above problems in the prior art, embodiments of the present disclosure provide a data query method, system, electronic device, storage medium, and program product.

[0004] A first aspect of the embodiments of the present disclosure provides a data query method, including:

[0005] Receiving a query request for business data in a first storage component of a business system;

[0006] Parsing the query request to obtain a first storage unit of the business data in the first storage component;

[0007] Querying a pre-set mapping table to obtain a second storage unit corresponding to the first storage unit; the second storage unit is a storage unit created correspondingly in a second storage component for storing business data;

[0008] Obtaining a query result corresponding to the query request from the second storage unit; wherein, the first storage component and the second storage component provide different data query functions respectively.

[0009] As a possible implementation manner of the first aspect, the process of pre-setting the mapping table includes:

[0010] Obtaining first metadata of business data from a first storage component of a business system, where the first metadata includes a first storage unit for storing business data in the first storage component;

[0011] Correspondingly creating a second storage unit for storing business data in the second storage component according to the first metadata;

[0012] Writing the corresponding relationship between the first storage unit and the second storage unit into the mapping table.

[0013] As a possible implementation of the first aspect, after setting up the mapping table, it further includes:

[0014] Create a write task based on the correspondence between the first storage unit and the second storage unit, where the write task is used to instruct the business system to write business data into the second storage component.

[0015] As a possible implementation of the first aspect, obtaining the first metadata of business data from the first storage component of the business system includes:

[0016] Obtain the operation instruction information of the user on the business data from the first storage component of the business system;

[0017] Parse the operation instruction information to obtain the first metadata of the business data.

[0018] As a possible implementation of the first aspect, parsing the operation instruction information to obtain the first metadata of the business data includes:

[0019] Parse the operation instruction information into an abstract syntax tree;

[0020] Traverse each node of the abstract syntax tree, and filter out the nodes whose data source is the first storage component and represent data writing;

[0021] Extract the first metadata of the business data from the filtered nodes.

[0022] As a possible implementation of the first aspect, according to the first metadata, correspondingly create a second storage unit for storing business data in the second storage component, including:

[0023] Generate second metadata for correspondingly creating the second storage unit according to the first metadata;

[0024] Based on the second metadata, correspondingly create a second storage unit for storing business data in the second storage component.

[0025] As a possible implementation of the first aspect, the business system adopts a Lambda architecture; the first storage component includes Kafka; the second storage component includes at least one of Paimon, MySQL, and StarRocks.

[0026] The second aspect of the embodiments of the present disclosure provides a data query device, including:

[0027] A receiving unit, configured to receive a query request for business data in the first storage component of the business system;

[0028] A parsing unit, configured to parse a query request to obtain a first storage unit of service data in a first storage component;

[0029] A query unit, configured to query a pre-set mapping table to obtain a second storage unit corresponding to the first storage unit; the second storage unit is a storage unit corresponding to be created in a second storage component for storing service data;

[0030] An obtaining unit, configured to obtain a query result corresponding to the query request from the second storage unit; wherein, the first storage component and the second storage component respectively provide different data query functions.

[0031] As a possible implementation manner of the second aspect, the above-mentioned apparatus further includes a setting unit, and the setting unit is configured to pre-set the mapping table, specifically including:

[0032] An obtaining subunit, configured to obtain first metadata of service data from a first storage component of a service system, and the first metadata includes a first storage unit for storing service data in the first storage component;

[0033] A creating subunit, configured to correspondingly create a second storage unit for storing service data in the second storage component according to the first metadata;

[0034] A writing subunit, configured to write the corresponding relationship between the first storage unit and the second storage unit into the mapping table.

[0035] As a possible implementation manner of the second aspect, the above-mentioned apparatus further includes a creating unit, and the creating unit is configured to:

[0036] After setting the mapping table, create a writing task based on the corresponding relationship between the first storage unit and the second storage unit, and the writing task is used to instruct the service system to write service data into the second storage component.

[0037] As a possible implementation manner of the second aspect, the obtaining subunit is configured to:

[0038] Obtain operation instruction information of a user on service data from a first storage component of a service system;

[0039] Parse the operation instruction information to obtain first metadata of service data.

[0040] As a possible implementation manner of the second aspect, the obtaining subunit is configured to:

[0041] Parse the operation instruction information into an abstract syntax tree;

[0042] Traverse each node of the abstract syntax tree, and filter out nodes whose data source is the first storage component and represent data writing;

[0043] Extract the first metadata of the service data from the selected nodes.

[0044] As a possible implementation of the second aspect, create a subunit for:

[0045] Generate the second metadata for correspondingly creating the second storage unit according to the first metadata;

[0046] Based on the second metadata, correspondingly create a second storage unit for storing service data in the second storage component.

[0047] As a possible implementation of the second aspect, the service system adopts a Lambda architecture; the first storage component includes Kafka; the second storage component includes at least one of Paimon, MySQL, and StarRocks.

[0048] A third aspect of the embodiments of the present disclosure provides a service system, including:

[0049] A first storage component and a second storage component for storing the service data of the service system;

[0050] A data query device as described in any one of the above second aspects, configured to receive a query request for the service data in the first storage component of the service system; parse the query request to obtain the first storage unit of the service data in the first storage component; query a pre-set mapping table to obtain the second storage unit corresponding to the first storage unit; the second storage unit is a storage unit correspondingly created in the second storage component for storing service data; obtain a query result corresponding to the query request from the second storage unit; wherein, the first storage component and the second storage component respectively provide different data query functions.

[0051] A fourth aspect of the embodiments of the present disclosure provides an electronic device, including:

[0052] A memory for storing a computer program product;

[0053] A processor for executing the computer program product stored in the memory, and when the computer program product is executed, implement the method according to any one of the above first aspects.

[0054] A fifth aspect of the embodiments of the present disclosure provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, implement the method according to any one of the above first aspects.

[0055] A sixth aspect of the present disclosure provides a computer program product, including computer program instructions, and when the computer program instructions are executed by a processor, implement the method according to any one of the above first aspects.

[0056] Based on the embodiments of the present disclosure, first, a query request for business data in the first storage component of the business system is received, and the second storage unit storing the business data in the second storage component is obtained by querying a pre-set mapping table. Based on the above solution, the second storage unit can be used as a redundant backup unit for the business data, and the main process of business processing is still based on the first storage component. Among them, the second storage component can provide a data query function different from that of the first storage component, can provide a data query ability that better meets the user's needs, and improve the data query efficiency.

[0057] The technical solution of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0058] The drawings forming a part of the specification depict the embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.

[0059] Referring to the accompanying drawings, the present disclosure can be more clearly understood according to the following detailed description, where:

[0060] Figure 1 It is a schematic diagram of the Lambda architecture in the related art;

[0061] Figure 2 It is a schematic diagram of data reading in an embodiment of the data query method of the present disclosure;

[0062] Figure 3 It is a flowchart of an embodiment of the data query method of the present disclosure;

[0063] Figure 4 It is a flowchart of an embodiment of the data query method of the present disclosure;

[0064] Figure 5 It is a schematic diagram of data writing in an embodiment of the data query method of the present disclosure;

[0065] Figure 6 It is a flowchart of an embodiment of the data query method of the present disclosure;

[0066] Figure 7 It is a schematic diagram of the lakehouse architecture in the related art;

[0067] Figure 8 It is a schematic structural diagram of an embodiment of the data query device of the present disclosure;

[0068] Figure 9 It is a schematic structural diagram of an embodiment of the data query device of the present disclosure;

[0069] Figure 10 It is a schematic structural diagram of an embodiment of the data query device of the present disclosure;

[0070] Figure 11 Block diagram of the electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0071] In order to accurately describe the technical content in the present disclosure and for accurate understanding of the present disclosure, the following explanatory notes or definitions are given to the terms used in this specification before the detailed implementation manners are described:

[0072] 1) Kafka: Kafka is a high-throughput distributed publish-subscribe messaging system that can process all action stream data of consumers in a website.

[0073] 2) Lambda architecture: It is a big data processing architecture specifically designed for processing large-scale data. The Lambda architecture aims to combine the advantages of stream processing and batch processing to provide comprehensive, accurate and low-latency data processing capabilities.

[0074] 3) Hive (data warehouse management system): Hive is a data warehouse tool based on Hadoop, mainly used for processing and analyzing large-scale structured data. The main functions of Hive include data extraction, transformation and loading, and it can store, query and analyze large-scale data stored in Hadoop. It is particularly suitable for scenarios that require processing a large amount of structured data, such as data analysis, data mining and complex queries, etc.

[0075] 4) StarRocks: StarRocks is an open-source new generation of extremely fast full-scenario MPP (Massively Parallel Processing) database. It adopts a new generation of elastic MPP architecture and can efficiently support various data analysis scenarios such as multi-dimensional analysis, real-time analysis, and high-concurrency analysis of large data volumes.

[0076] 5) Point Lookup: In a real-time data lake, Point Lookup is a query operation that quickly finds specific records in a dimension table through one or a small number of primary keys or index fields.

[0077] 6) Kafka2Kafka: It mainly involves data synchronization or migration between multiple Kafka clusters, referring to the process of real-time transmitting data in a topic in one Kafka cluster to another topic in another Kafka cluster. This operation is usually used in scenarios such as data migration, data backup or load balancing.

[0078] 7) Kafka2StarRocks: It refers to the process of importing data from Kafka into StarRocks, and it is a tool or solution for importing Kafka data into StarRocks.

[0079] 8) Apache Paimon: It is a streaming data lake storage technology aiming to provide high-throughput, low-latency data ingestion, streaming subscription, and real-time query capabilities.

[0080] 9) Apache Flink: It is an open-source stream processing and batch processing framework mainly used for distributed computing of real-time data streams and large-scale data batch processing. It is developed by the Apache Software Foundation, supports the processing of unbounded and bounded data streams, and has characteristics such as high throughput, low latency, and fault tolerance.

[0081] 10) Apache Calcite: It is a dynamic data management framework mainly used for SQL (Structured Query Language) parsing, query optimization, and execution. It can parse various SQL statements into an abstract syntax tree (AST), optimize the query plan by operating on the AST, and finally generate code that can be executed on a specific data processing engine.

[0082] 11) MySQL: It is a relational database management system. Relational databases store data in different tables instead of putting all data in a large warehouse, which increases speed and flexibility.

[0083] First, the existing methods will be introduced below, and then the technical solutions of the present disclosure will be introduced in detail.

[0084] Generally, the storage components in a business system (such as a business system adopting the Lambda architecture) may only support a certain specific data query method and can only provide limited data query functions. For example, the data in a message queue mainly based on Kafka (Kafka) has a validity period and only supports querying through Offset (offset), and cannot be filtered according to field content. Therefore, when data quality problems occur, it is difficult to locate the problems, and the troubleshooting efficiency is very low, making it difficult to meet the user's data query needs.

[0085] See Figure 1, taking a business system adopting the Lambda architecture as an example, to solve the above problems, in the related technology, in addition to the real-time link, an offline link is introduced. The offline link also accesses from the business database and periodically imports data into the offline data warehouse. For example, Hive (a data warehouse management system) is used to store offline data, and Hive incremental data is obtained in a hourly scheduling manner, and the Hive incremental data is merged with the Hive stored data. The offline link mainly has two functions: one is to default that the offline data is accurate and periodically correct the real-time data through the offline data; the other is to locate the problem data by querying the offline table when the business conducts data troubleshooting. In Figure 1 , Kafka and StarRocks are used to store real-time data, and Flink is used to process online data. According to the data troubleshooting needs, the offline data and real-time data can be queried respectively.

[0086] However, in terms of data troubleshooting, this Lambda architecture also has the following obvious defects:

[0087] 1. The offline data warehouse has a scheduling cycle of hours and days, and data troubleshooting has a lag. For example, if you want to locate the data at time H, you need to wait until time H+2, and you can query after the new partition is created.

[0088] 2. Some types of real-time tables support updates, while the offline data warehouse does not. Therefore, the offline part needs to maintain two sets of logics for incremental and stored data, and perform periodic merging of incremental and stored data. The offline and real-time links under this Lambda architecture may require two different implementation logics and are based on different computing engines, resulting in high operation and maintenance costs.

[0089] In summary, the existing technology has the following defects: limited data query function, poor timeliness, low data query efficiency, and it is difficult to meet the data query needs of users.

[0090] Based on the technical problems existing in the above-mentioned prior art, the embodiment of the present disclosure provides a data query method, which first receives a query request for business data in a first storage component of a business system, and obtains a second storage unit storing business data in a second storage component by querying a pre-set mapping table. Based on the above scheme, the second storage unit can be used as a redundant backup unit for business data, the main process of business processing is still based on the first storage component, and the business data generated in the first storage component is backed up in real time to the second storage component during the business processing. Among them, the second storage component can provide a data query function different from the first storage component, while ensuring timeliness, providing a data query capability that better meets user needs, and improving data query efficiency, thereby solving the technical problems mentioned in the prior art that the data query function is limited, the timeliness is poor, the data query efficiency is low, and it is difficult to meet the data query needs of users.

[0091] In a specific example, the first storage component used to store business data in the business system, such as Kafka, cannot be filtered according to the field content, does not support the point query function, and cannot meet the user's data query needs. In the disclosed embodiment, in order to implement the above-mentioned data query method, it is necessary to perform a double write operation on the business data in the message queue of the business system, that is, in the normal business processing link, the business data is written to the first storage component, and outside the normal business processing link, an additional second storage component that supports point query is written, for example, Paimon can be used as the second storage component. At the same time, a set of microservices is developed to automatically create double write tasks and pre-set a mapping table, which includes the corresponding relationship between the business data double-written to the storage units in the first storage component and the second storage component. When the user initiates a query request for business data, the microservice converts the user's query message queue request into a query request for the second storage component by querying the underlying mapping table, and forwards the user's query request to the second storage component. After obtaining the query result, the query result is returned step by step, so as to achieve user insensitivity.

[0092] In the above business system, the tasks created in the business process of data processing may include real-time computing tasks and dual-write tasks. Among them, the real-time computing task can be a real-time computing task developed by the business party, and the Kafka data in the message queue generated by the real-time computing task is written to the first storage component, for example Figure 1 The dual-write task is used to persist the message queue. It is automatically created by the microservice and writes the Kafka data to the second storage component Paimon while the real-time computing task of the business side generates Kafka data.

[0093] In addition, in the above business system, a real-time computing platform can be set up. The real-time computing platform directly interacts with users. Users create data processing tasks and query data through the real-time computing platform, and the underlying microservices are hidden from users. Figure 2 It is a schematic diagram of data reading for an embodiment of the data query method of the present disclosure. Refer to Figure 2 As shown, the user sends a query request for business data to the real-time computing platform. The real-time computing platform forwards the query request to the microservice, receives the query result returned by the microservice, and then returns the query result to the user.

[0094] Figure 3 It is a flowchart of an embodiment of the data query method of the present disclosure. The embodiments of the present disclosure can be applied to a business system adopting the Lambda architecture or other architecture business systems. The embodiments of the present disclosure are described by taking the application to a business system adopting the Lambda architecture as an example. As Figure 3 shown, this method can be applied to microservices and specifically may include:

[0095] Step S110, receive a query request for business data in the first storage component of the business system.

[0096] Refer to Figure 2 and Figure 3 For the example of, Kafka is used as the first storage component, and the business data generated by the real-time computing task is written into the first storage component. When the user sends a query request for the business data in the first storage component to the real-time computing platform, the real-time computing platform forwards the query request to the microservice. The microservice receives the query request for the business data in the first storage component of the business system.

[0097] Step S120, parse the query request to obtain the first storage unit of the business data in the first storage component.

[0098] The first storage component includes multiple first storage units for storing business data, and each first storage unit can specifically use a data table as the data storage form. In the Figure 2 example of, a Kafka table is used as the first storage unit. The query request may include the name of the data table that the user wants to query. For example, the name of the data table is Kafka1. The query request may also include query fields and query ranges. For example, query fields and corresponding filtering conditions can be given. In addition, the query request may also include information about the server where the data table to be queried is located. The microservice parses the received query request to obtain the first storage unit of the business data in the first storage component. In the Figure 2 example of, the microservice parses the query request to obtain the information of the Kafka table where the business data is located.

[0099] Step S130, query the pre-set mapping table to obtain the second storage unit corresponding to the first storage unit; the second storage unit is a storage unit created correspondingly in the second storage component for storing service data.

[0100] When performing a dual-write operation on service data, first, the microservice automatically creates a dual-write task to back up the service data generated by the real-time computing task to the first storage component at the same time, and writes the corresponding relationship between the service data and the storage units in the first storage component and the second storage component into the mapping table. In one example, in the service link, the service data is stored in the first storage unit in the first storage component, and the first storage unit is a data table named Kafka1. When performing the dual-write operation, the microservice creates a second storage unit for storing the service data in the second storage component, uses Paimon as the second storage component, and the second storage unit is a data table named Paimon1. The microservice creates a dual-write task to dual-write the service data to the storage unit Kafka1 in the first storage component and the storage unit Paimon1 in the second storage component respectively. At the same time, the microservice maintains the pre-set mapping table and writes the corresponding relationship between the first storage unit Kafka1 and the second storage unit Paimon1 into the mapping table.

[0101] In step S120, the microservice parses the query request to obtain the first storage unit of the service data in the first storage component, that is, it is obtained that the service data is stored in the data table Kafka1 in the first storage component. In step S130, the microservice queries the mapping table to obtain the second storage unit corresponding to the first storage unit, that is, it is obtained that the corresponding second storage unit of the data table Kafka1 in the first storage component in the second storage component is the data table Paimon1.

[0102] Step S140, obtain the query result corresponding to the query request from the second storage unit; wherein, the first storage component and the second storage component respectively provide different data query functions.

[0103] Based on the fact that in step S130, it is queried in the mapping table that the service data is stored in the data table Paimon1 in the second storage component, in step S140, the microservice queries the second storage unit, obtains the query result corresponding to the query request from the data table Paimon1, and then returns the query result to the user level by level. In the above example, the first storage component cannot perform filtering according to field content and does not support point query function. By dual-writing the service data to the second storage component that supports querying according to content, the query needs of the user can be met by querying the second storage component.

[0104] In Figure 2In the example, the data reading process includes the following steps:

[0105] ① The user sends a query request to the real-time computing platform;

[0106] ② The real-time computing platform forwards the query request to the microservice;

[0107] ③ The microservice queries the mapping table and requests to obtain the Paimon table information;

[0108] ④ The microservice receives the returned Paimon table information;

[0109] ⑤ The microservice queries the Paimon table;

[0110] ⑥ The microservice receives the query result returned by the Paimon table;

[0111] ⑦ The microservice returns the query result to the real-time computing platform;

[0112] ⑧ The real-time computing platform returns the query result to the user.

[0113] Based on the embodiments of the present disclosure, first, a query request for business data in the first storage component of the business system is received, and the second storage unit storing the business data in the second storage component is obtained by querying a pre-set mapping table. Based on the above solution, the second storage unit can be used as a redundant backup unit for business data. The main process of business processing is still based on the first storage component, and the business data generated in the first storage component is real-time backed up to the second storage component during the business processing. Among them, the second storage component can provide a data query function different from that of the first storage component, providing a data query ability that better meets user needs while ensuring timeliness and improving data query efficiency.

[0114] As Figure 4 shown, in one implementation, the process of pre-setting the mapping table includes:

[0115] Step S210, obtaining the first metadata of the business data from the first storage component of the business system, where the first metadata includes the first storage unit for storing the business data in the first storage component.

[0116] Figure 5 This is a schematic diagram of data writing for an embodiment of the data query method of the present disclosure. Refer to Figure 4 and Figure 5, in the normal business processing link of the business system, after the data in the data table Kafka2 is processed by Flink, the obtained new business data is stored in the data table Kafka3, that is, Kafka3 is the first storage unit of the first storage component where the business data is located. At the same time, the information of the first metadata (Kafka DDL) of the business data is sent to the microservice. The first metadata includes the first storage unit in the first storage component for storing business data, that is, the name of the data table Kafka3. The first metadata also includes data description information such as the definition of the table Kafka3, table name, data format, parsing method, fields, primary key, etc. of the data table. The microservice receives this information and obtains the first metadata of the business data.

[0117] Step S220, according to the first metadata, correspondingly create a second storage unit for storing business data in the second storage component.

[0118] The microservice parses the first metadata Kafka DDL to obtain the first storage unit of the business data and data description information such as the definition of the table, table name, data format, parsing method, fields, primary key, etc. The microservice correspondingly creates a second storage unit for storing business data in the second storage component according to the data description information. See Figure 5 , specifically, a second storage unit for storing business data can be correspondingly created in the second storage component Paimon, that is, the data table Paimon3.

[0119] Step S230, write the corresponding relationship between the first storage unit and the second storage unit into the mapping table.

[0120] After creating the second storage unit, the microservice writes the corresponding relationship between the first storage unit and the second storage unit into the mapping table. See Figure 5 , specifically, write the corresponding relationship between the data table Kafka3 in the first storage component and the data table Paimon3 in the second storage component into the mapping table.

[0121] See Figure 5 , read the source Kafka table structure and the written Paimon table structure are stored in the mapping table. The mapping table structure is as follows:

[0122] id source source_config sink sink_config 1 Kafka_table_1 {"type":"Kafka"…} Paimon_table_1 {"type":"Paimon"…} 2 Kafka_table_2 {"type":"Kafka"…} Paimon_table_2 {"type":"Paimon"…}

[0123] Different types of metadata information are stored in config, and their content meanings are as follows:

[0124] 1) type: data source type, including Kafka, Paimon, and can be expanded to other types such as MySQL, StarRocks, etc. The mapping table can store data from various data sources, and data backup mapping can be done between various databases, that is, the mapping between source and sink. The source in the mapping table represents the data table that reads data, and the sink represents the data table that writes data. For example, if the data in Kafka is written to Paimon, the source_config describes some metadata of Kafka, such as the name of the Kafka table where the data is located; sink_config describes some metadata of Paimon, such as the name of the Paimon table where the data is located. The config can include the storage address of the data table, field information, etc. The table names "Kafka_table_1, Kafka_table_2" in the above mapping table correspond to Figure 2 and Figure 5 "Kafka1, Kafka2" in the

[0125] Since various data sources have their own limitations, each data source can only provide limited data query functions. In the business processing link, business data is stored in the first data source, and dual write operations are performed on the business data. Backup data is stored in the second data source. The storage location of the backup data can be queried through the mapping table, and data can be queried through the second data source. The second data source can provide query functions different from the first data source, thereby removing the limitations of the data query function.

[0126] 2) columns: The column information of the table, in the format of List, describes the field information of the table.

[0127] 3) meta: Metadata information of the table. Metadata includes the storage address, which can uniquely locate a data source. Different types of data sources may store different metadata information. For example, Kafka needs to specify a broker (server) to locate a data source, that is, specify a server address; and a data source name, that is, the name of the data table. Paimon needs to specify a path to locate a data source.

[0128] In the embodiment of the present disclosure, a mapping table is pre-set to establish a correspondence between the storage units in the first storage component and the second storage component that store the same data. Therefore, when a user initiates a query request for the business data of the first storage component, the second storage unit in the second storage component can be obtained by querying the mapping table. Based on the query function provided by the second storage component that is different from that of the first storage component, the user's query needs can be better met.

[0129] likeFigure 6 As shown, in one embodiment, after setting up the mapping table, it further includes:

[0130] Step S240, creating a write task based on the correspondence between the first storage unit and the second storage unit, where the write task is used to instruct the business system to write business data into the second storage component.

[0131] Refer again to Figure 5 , the microservice automatically creates a write task based on the correspondence between the first storage unit and the second storage unit, instructing the business system to write the data in the data table Kafka3 of the first storage component into the data table Paimon3 of the second storage component. After the microservice generates a write task from Kafka to Paimon (Kafka2Paimon), it submits the generated write task to the big data cluster device running the task to start data synchronization.

[0132] In the embodiments of the present disclosure, by performing dual-write operations on business data, the second storage unit is used as a redundant backup unit for business data, and a data query capability that better meets user needs is provided based on the query function provided by the second storage component that is different from the first storage component, improving data query efficiency.

[0133] In one embodiment, step S210, obtaining the first metadata of business data from the first storage component of the business system, includes:

[0134] Obtaining the operation instruction information of the user on the business data from the first storage component of the business system;

[0135] Parsing the operation instruction information to obtain the first metadata of the business data.

[0136] Refer to Figure 5For example, in a business system, real-time computing tasks can generate a business processing flow from Kafka1 to Kafka2 and then from Kafka2 to Kafka3, and this business processing flow can be created based on the user's SQL (Structured Query Language) text. The user's SQL text includes operation instruction information for business data. For example, the operations can include adding, deleting, updating, and filtering data, etc. Specifically, Flink is used to process the data in the upstream table and write the processing results to the downstream table. For example, in the business processing flow from Kafka2 to Kafka3, Kafka2 is the upstream table and Kafka3 is the downstream table. The data in Kafka2 is processed and the processing results are written to Kafka3, and at the same time, the user's SQL text is sent to the microservice. The operation instruction information for business data also includes the first metadata of the business data. The microservice parses the operation instruction information in the SQL text and can obtain the first metadata of the business data.

[0137] In the embodiments of the present disclosure, the first storage unit for storing business data in the first storage component can be obtained through the first metadata. On this basis, a dual-write task is created and a dual-write operation is performed on the business data. The second storage unit is used as the redundant backup unit for the business data, and a data query capability that better meets the user's needs is provided based on the query function provided by the second storage component different from the first storage component, thereby improving the data query efficiency.

[0138] In one implementation, parsing the operation instruction information to obtain the first metadata of the business data includes:

[0139] Parsing the operation instruction information into an abstract syntax tree;

[0140] Traversing each node of the abstract syntax tree and filtering out the nodes whose data source is the first storage component and that represent data writing;

[0141] Extracting the first metadata of the business data from the filtered nodes.

[0142] The microservice parses the operation instruction information in the user's SQL text into an Abstract Syntax Tree (AST) through Apache Calcite. Among them, the abstract syntax tree is a tree-like data structure used to represent the structure of an SQL query, which is a way to transform an SQL statement into a tree-like structure. The data processing logic represented in the abstract syntax tree includes: source (upstream, data reading), sink (downstream, data writing), and transformation operations. According to the SQL text, the business system performs transformation operations on the data in the upstream table source and writes the processing results to the downstream table sink. The transformation operations can include extracting fields, performing addition, subtraction, multiplication, and division operations, or performing other processing, concatenation, etc. The data processing logic represented by the SQL text is described in the syntax tree, specifically including which storage unit to read data from, how to process the data after reading, and which storage unit to write the processed data to.

[0143] After the microservice parses the operation instruction information into an abstract syntax tree, it traverses each node in the syntax tree and filters out the nodes of the sink type with the data source of the Kafka type. Among them, the sink type represents the node where data is written, and the Kafka type represents that the data source is the first storage component. That is to say, the microservice filters out the nodes in the syntax tree with the data source of the first storage component and representing data writing. Then, the first metadata (Schema) of the business data is extracted from the filtered nodes.

[0144] In the embodiments of the present disclosure, by traversing and filtering the abstract syntax tree nodes, this tree-like structure can perform query optimization more effectively. It can more intuitively represent the organizational structure and logical relationship of the query, improve the efficiency of query processing, and thus improve the system performance and timeliness.

[0145] In one implementation, step S220, creating a second storage unit for storing business data in the second storage component according to the first metadata, includes:

[0146] Generating second metadata for correspondingly creating the second storage unit according to the first metadata;

[0147] Based on the second metadata, correspondingly creating a second storage unit for storing business data in the second storage component.

[0148] The first metadata includes data description information such as the definition of a table, data format, fields, primary key, etc. For example, if Kafka is used as the first storage component and Paimon is used as the second storage component, the corresponding first metadata and second metadata have different data formats respectively. The microservice generates the second metadata for creating the corresponding second storage unit based on the data description information in the first metadata. Then, based on the second metadata, a second storage unit for storing business data is correspondingly created in the second storage component, that is, a Paimon table corresponding to the Kafka table is created.

[0149] In Figure 5 the example of, the data writing process includes the following steps:

[0150] ① The first metadata Kafka DDL is passed from the business processing link to the microservice;

[0151] ② The microservice parses the Kafka DDL and generates the DDL of the Paimon table, that is, the second metadata;

[0152] ③ The microservice creates a Paimon table and a write task from Kafka to Paimon (Kafka2Paimon).

[0153] ④ Write the corresponding relationship from Kafka to Paimon into the mapping table;

[0154] ⑤ Perform real-time data writing from Kafka to Paimon.

[0155] To avoid intrusion into the original task, the task of synchronizing Kafka data to Paimon in step ⑤ can be performed asynchronously, that is, this task is an independent task and is independently submitted to the big data cluster device running the task.

[0156] In the embodiments of the present disclosure, by correspondingly creating a second storage unit for storing business data in the second storage component, the dual-write operation of business data is realized, so that the second storage unit is used as a redundant backup unit for business data, and a data query capability that better meets user needs is provided based on the query function provided by the second storage component, which is different from that of the first storage component, and the data query efficiency is improved.

[0157] In one implementation, the business system adopts a Lambda architecture; the first storage component includes Kafka; the second storage component includes at least one of Paimon, MySQL, and StarRocks.

[0158] In the related art, for the application of a streaming data lake, the architecture of "Flink + lake warehouse" is used to replace Figure 1 the architecture of "Flink + Kafka" in. The replaced architecture is asFigure 7 As shown, where ODS (Operation Data Store), DWD (Data Warehouse Detail), and DWS (Data Warehouse Service) are different levels in the data warehouse, representing the operation data storage layer, the detailed data layer, and the service data layer respectively; "LakeHouse" represents the integration of lake and warehouse; "Changelog" represents the change log; and "MySQL" is a relational database management system. Figure 7 Although the replaced architecture in [description] can achieve point queries on Kafka data, the lake storage is based on distributed storage such as HDFS (Hadoop Distributed File System), and the data reading and writing are slower than Kafka. The single-task latency is at the minute level, and the overall link latency is at the level of several minutes.

[0159] Compared with Figure 1 the Lambda architecture that combines real-time and offline links, the embodiment of the present disclosure adopts an architecture that combines two real-time links with the help of a cutting-edge streaming data lake, which can fundamentally solve the technical problems of limited data query function and poor timeliness. In the embodiment of the present disclosure, for the real-time data warehouse based on Flink, Apache Paimon is selected as the lake-warehouse product in the solution. See Figure 2 and Figure 5 As shown, the architecture combined with Flink and Paimon is used to replace Figure 1 the Hive2Hive offline data warehouse in [description]. The lake-warehouse is used as an auxiliary link, and the main business process is still based on Kafka. The first storage component is Kafka; the first storage component is Paimon, which provides data query capabilities while ensuring timeliness. This solution has the following beneficial technical effects:

[0160] 1. In terms of timeliness, all calculations are based on the Flink real-time computing engine. The main process is still based on the architecture combined with Flink and Kafka, and the main business output remains at the second level. At the same time, the data synchronization from Kafka to Paimon is also based on Flink, fundamentally avoiding hourly scheduling, and reducing the timeliness of data point queries from H+2 to the minute level.

[0161] 2. In terms of cost, the dual-write link can be based on the same real-time computing engine, avoiding two different engines for real-time and offline links, unifying the syntax, and effectively reducing the operation and maintenance costs of developers.

[0162] 3. Asynchronous writing with dynamic introduction of new tasks. That is, when business data is written from Kafka to Kafka during runtime, a dual-writing task is automatically created in the background, and a writing task from Kafka to Paimon is automatically generated. There is no intrusion into the original task and no perception by users.

[0163] In addition, in other implementation manners of the embodiments of the present disclosure, the first storage component includes Kafka; the second storage component may further include at least one of MySQL and StarRocks. The second storage component can provide a data query function different from that of the first storage component, providing a data query ability that better meets user needs while ensuring timeliness and improving data query efficiency.

[0164] As Figure 8 shown, the present disclosure also provides an embodiment of a corresponding data query device. For the beneficial effects or technical problems solved by this device, reference can be made to the descriptions in the methods corresponding to each device or the description in the summary of the invention, which will not be elaborated here one by one.

[0165] In the embodiment of this data query device, the device includes:

[0166] A receiving unit 100, configured to receive a query request for business data in the first storage component of the business system;

[0167] An analysis unit 200, configured to analyze the query request to obtain a first storage unit of the business data in the first storage component;

[0168] A query unit 300, configured to query a pre-set mapping table to obtain a second storage unit corresponding to the first storage unit; the second storage unit is a storage unit created in the second storage component for storing business data;

[0169] An obtaining unit 400, configured to obtain a query result corresponding to the query request from the second storage unit; wherein, the first storage component and the second storage component respectively provide different data query functions.

[0170] As Figure 9 and Figure 10 shown, in one implementation manner, the above device further includes a setting unit 500, and the setting unit 500 is configured to pre-set the mapping table, specifically including:

[0171] An obtaining subunit 510, configured to obtain first metadata of business data from the first storage component of the business system, and the first metadata includes a first storage unit for storing business data in the first storage component;

[0172] Create a sub-unit 520, which is used to correspondingly create a second storage unit for storing business data in the second storage component according to the first metadata;

[0173] A write sub-unit 530, which is used to write the corresponding relationship between the first storage unit and the second storage unit into the mapping table.

[0174] As Figure 9 shown, in one implementation, the above device further includes a creation unit 600, and the creation unit 600 is used for:

[0175] After setting the mapping table, create a write task based on the corresponding relationship between the first storage unit and the second storage unit, and the write task is used to instruct the business system to write business data into the second storage component.

[0176] In one implementation, the acquisition sub-unit 510 is used for:

[0177] Obtain the operation instruction information of the business data from the first storage component of the business system;

[0178] Parse the operation instruction information to obtain the first metadata of the business data.

[0179] In one implementation, the acquisition sub-unit 510 is used for:

[0180] Parse the operation instruction information into an abstract syntax tree;

[0181] Traverse each node of the abstract syntax tree, and filter out the nodes whose data source is the first storage component and represent data writing;

[0182] Extract the first metadata of the business data from the filtered nodes.

[0183] In one implementation, the creation sub-unit 520 is used for:

[0184] Generate second metadata for correspondingly creating a second storage unit according to the first metadata;

[0185] Based on the second metadata, correspondingly create a second storage unit for storing business data in the second storage component.

[0186] In one implementation, the business system adopts a Lambda architecture; the first storage component includes Kafka; the second storage component includes at least one of Paimon, MySQL, and StarRocks.

[0187] The data query device in the embodiments of the present disclosure corresponds to the above-mentioned data query method embodiments of the present disclosure in terms of specific implementation and beneficial technical effects, and the relevant content can be referred to each other, and will not be elaborated here.

[0188] The present disclosure also provides a corresponding embodiment of a service system. For the beneficial effects or technical problems solved by the service system, reference may be made to the descriptions in the above data query device and the methods corresponding to each device, or the descriptions in the summary of the invention. Details will not be repeated herein.

[0189] In the embodiment of the service system, the service system includes:

[0190] A first storage component and a second storage component for storing the service data of the service system;

[0191] The above data query device is configured to receive a query request for the service data in the first storage component of the service system; parse the query request to obtain a first storage unit of the service data in the first storage component; query a pre-set mapping table to obtain a second storage unit corresponding to the first storage unit; the second storage unit is a storage unit created correspondingly in the second storage component for storing the service data; obtain a query result corresponding to the query request from the second storage unit; wherein, the first storage component and the second storage component respectively provide different data query functions.

[0192] The exemplary embodiments of the data query method, data query device and service system of the present disclosure correspond to each other, and can be mutually referred to and cited in the corresponding implementation manners, details will not be repeated.

[0193] Next, with reference to Figure 11 to describe an electronic device according to an embodiment of the present disclosure. The electronic device may be any one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device can communicate with the first device and the second device to receive the input signals collected from them.

[0194] Figure 11 The block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.

[0195] As Figure 11 shown, the electronic device includes one or more processors and a memory.

[0196] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0197] The memory can store one or more computer program products. The memory can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products can be stored on the computer-readable storage media, and the processor can run the computer program products to implement the data query methods of the various embodiments of the present disclosure above and / or other desired functions.

[0198] In one example, the electronic device may further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0199] In addition, the input device may further include, for example, a keyboard, a mouse, and so on.

[0200] The output device can output various information to the outside, including the determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.

[0201] Of course, for simplicity, Figure 11 only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.

[0202] In addition to the above methods and devices, the embodiments of the present disclosure may also be a computer program product, which includes computer program instructions that cause the processor to execute the steps in the data query methods according to various embodiments of the present disclosure described in the above part of this specification when the computer program instructions are run by the processor.

[0203] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0204] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the data query method according to various embodiments of the present disclosure described in the foregoing part of this specification.

[0205] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0206] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-described specific details are only for illustrative purposes and for ease of understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0207] Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system embodiments, since they basically correspond to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0208] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0209] The methods and apparatuses of the present disclosure may be implemented in many ways. For example, the methods and apparatuses of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present disclosure. Thus, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.

[0210] It should also be noted that in the apparatuses, devices, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0211] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0212] The above description has been presented for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and subcombinations thereof.

Claims

1. A data query method, characterized in that: include: Receiving a query request for business data in a first storage component of a business system; Parsing the query request to obtain a first storage unit of the business data in the first storage component; Querying a preset mapping table to obtain a second storage unit corresponding to the first storage unit; the second storage unit is a storage unit correspondingly created in the second storage component for storing the business data; Obtaining a query result corresponding to the query request from the second storage unit; wherein the first storage component and the second storage component respectively provide different data query functions.

2. The method according to claim 1, characterized in that The process of presetting the mapping table includes: Acquire first metadata of business data from a first storage component of the business system, where the first metadata includes a first storage unit in the first storage component for storing the business data; According to the first metadata, a second storage unit for storing the business data is correspondingly created in the second storage component; The corresponding relationship between the first storage unit and the second storage unit is written into the mapping table.

3. The method according to claim 2, characterized in that After setting the mapping table, the method further includes: A write task is created based on the correspondence between the first storage unit and the second storage unit, where the write task is used to instruct the business system to write the business data into the second storage component.

4. The method according to claim 2 or 3, characterized in that: The acquiring first metadata of the business data from the first storage component of the business system includes: Acquire user operation instruction information on business data from the first storage component of the business system; The operation instruction information is parsed to obtain first metadata of the business data.

5. The method according to claim 4, characterized in that The parsing of the operation instruction information to obtain the first metadata of the business data includes: Parsing the operation instruction information into an abstract syntax tree; Traversing each node of the abstract syntax tree, and filtering out nodes whose data source is the first storage component and which represent data writing; The first metadata of the business data is extracted from the filtered nodes.

6. The method according to claim 2 or 3, characterized in that: The step of creating a second storage unit for storing the business data in a second storage component according to the first metadata includes: generating, according to the first metadata, second metadata for correspondingly creating the second storage unit; Based on the second metadata, a second storage unit for storing the business data is correspondingly created in the second storage component.

7. The method according to any one of claims 1 to 3, characterized in that The business system adopts Lambda architecture; the first storage component includes Kafka; the second storage component includes at least one of Paimon, MySQL, and StarRocks.

8. A business system, characterized in that: include: A first storage component and a second storage component are used to store business data of the business system; A data query device, used to receive a query request for business data in a first storage component of a business system; Parsing the query request to obtain a first storage unit of the business data in the first storage component; Querying a preset mapping table to obtain a second storage unit corresponding to the first storage unit; the second storage unit is a storage unit correspondingly created in the second storage component for storing the business data; Obtaining a query result corresponding to the query request from the second storage unit; wherein the first storage component and the second storage component respectively provide different data query functions.

9. An electronic device, characterized in that: include: a memory for storing a computer program product; A processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, it implements the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.

11. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.