Data processing method and apparatus, electronic device and storage medium
By using the intermediate index layer information for real-time query and sorting during data writing to disk, the problem that data cannot provide query services during DDL reorganization is solved, and the response speed of database business processing is improved and timeout problems are avoided.
Patent Information
- Application Number
- PCT/CN2024/128079
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-10-29
- Publication Date
- 2025-06-19
AI Technical Summary
In the process of reorganizing data using DDL, the relevant data cannot provide query services to the outside world, resulting in the impact of the database's business processing and business request timeout.
A data processing method is proposed. In the process of writing target data in memory to disk, in response to receiving a query instruction for target data, according to the query instruction and the intermediate index layer information of at least one first local data and at least one second local data, respectively, query the query data and sort the query data to provide real-time query services.
It realizes the real-time query service when data is written to disk by memory, ensures the accuracy of data query results, and avoids the business processing timeout caused by data reorganization.
Smart Images

Figure CN2024128079_19062025_PF_FP_ABST
Abstract
Description
Data processing method and device, electronic device and storage medium Technical Field
[0001] One or more embodiments of the present specification relate to the field of database technology, and in particular, to a data processing method and apparatus, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of the internet and information technology, data generation is exploding, placing increasing demands on databases and their management. Data processing requires the use of DML (Data Manipulation Language) to manipulate tables, such as adding, deleting, querying, and modifying data. Data processing also requires the use of DDL (Data Definition Language) to restructure data, such as creating new tables, deleting columns, and changing column types.
[0003] In related technologies, when using DDL to reorganize data, the relevant data cannot provide external query services, which affects the business processing of the database and causes business problems such as business request timeouts.
[0004] Summary of the Invention
[0005] According to a first aspect of one or more embodiments of the present specification, a data processing method is proposed, the method comprising: in a process of writing target data in a memory to a disk, in response to receiving a query instruction for the target data, querying data in the at least one first local data and the at least one second local data according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively, and sorting the queried data as a data query result; wherein, the first local data includes part of the target data stored on the disk, the second local data includes part of the target data stored in the memory, the target data is stored in a columnar form, and each column group of the target data corresponds to a second local data.
[0006] In one embodiment of the present specification, the query instruction includes a query range and a query condition; the query instruction, as well as the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively querying data in the at least one first local data and the at least one second local data, and sorting the queried data as a data query result, includes: according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively querying the data within the query range in the at least one first local data and the at least one second local data, and sorting the queried data as the data to be queried; querying the data to be queried according to the query condition to determine the data query result.
[0007] In one embodiment of the present specification, the query range includes a column query range of at least one column group; according to the query range, and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, the data within the query range is queried in the at least one first local data and the at least one second local data respectively, and the queried data is sorted as the data to be queried, including: for the column group to which each column query range in the query range belongs, according to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, and the column query range of the column group, determining the first row offset range corresponding to the column query range of the column group; determining the second row offset range according to the first row offset range corresponding to each column query range in the query range; according to the second row offset range, and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, the data within the second row offset range is queried in the at least one first local data and the at least one second local data respectively, and the queried data is sorted as the data to be queried.
[0008] In one embodiment of the present specification, determining the first row offset range corresponding to the column query range of the column group based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, and the column query range of the column group includes: querying the first local data and the second local data of the column group for data within a third row offset range within the query range based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, respectively, and sorting the queried data according to the row offset as the column query result of the column group; filtering the column query result of the column group based on the column query range of the column group to obtain the first row offset range corresponding to the column query range of the column group.
[0009] In one embodiment of the present specification, the query range includes a primary key range; the method further includes: querying the data within the primary key range in the first local data and the second local data of the primary key column respectively based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the primary key column, and sorting the queried data according to row offset as the primary key query result; determining the second row offset range within the query range based on the row offset of the data in the primary key query result.
[0010] In one embodiment of the present specification, the data within the first local data and the second local data are stored in the form of data blocks, and each database stores partial data; the data within the query range is queried in the at least one first local data and the at least one second local data according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and the queried data is sorted as the data to be queried, including: according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, the data blocks to which the data within the query range belong are queried in the at least one first local data and the at least one second local data, and the queried data blocks are sorted as the data to be queried.
[0011] In one embodiment of the present specification, the data blocks include macroblocks and microblocks.
[0012] In one embodiment of the present specification, the target data includes: data generated in the memory according to a data definition language DDL.
[0013] In one embodiment of the present specification, different column groups of the target data correspond to the same first partial data; or different column groups of the target data correspond to different first partial data.
[0014] According to a second aspect of one or more embodiments of the present specification, a data processing device is proposed, comprising: a real-time query module for, in the process of writing target data in memory to disk, responding to receiving a query instruction for the target data, querying data in the at least one first local data and the at least one second local data according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively, and sorting the queried data as a data query result; wherein, the first local data includes part of the target data stored on disk, the second local data includes part of the target data stored in memory, the target data is stored in column form, and each column group of the target data corresponds to one second local data.
[0015] In one embodiment of the present specification, the query instruction includes a query range and a query condition; the real-time query module is used to: query the data within the query range in the at least one first local data and the at least one second local data according to the intermediate index layer information of the at least one first local data and the intermediate index layer information of the at least one second local data, and sort the queried data as the data to be queried; query the data to be queried according to the query condition to determine the data query result.
[0016] In one embodiment of the present specification, the query range includes a column query range of at least one column group; the real-time query module is used to query the data within the query range in the at least one first local data and the at least one second local data according to the query range, the intermediate index layer information of at least one first local data, and the intermediate index layer information of at least one second local data, and sort the queried data as the data to be queried, and is used to: for the column group to which each column query range in the query range belongs, determine the first row offset range corresponding to the column query range of the column group according to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, and the column query range of the column group; determine the second row offset range according to the first row offset range corresponding to each column query range in the query range; query the data within the second row offset range in the at least one first local data and the at least one second local data according to the second row offset range, the intermediate index layer information of at least one first local data, and the intermediate index layer information of at least one second local data, and sort the queried data as the data to be queried.
[0017] In one embodiment of the present specification, the real-time query module is used to determine the first row offset range corresponding to the column query range of the column group based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, and the column query range of the column group. It is used to: query the first local data and the second local data of the column group for data within the third row offset range of the query range according to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, respectively, and sort the queried data according to the row offset as the column query result of the column group; filter the column query result of the column group according to the column query range of the column group to obtain the first row offset range corresponding to the column query range of the column group.
[0018] In one embodiment of the present specification, the query range includes a primary key range; the device also includes a primary key module, which is used to: query the data within the primary key range in the first local data and the second local data of the primary key column according to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the primary key column, and sort the queried data according to the row offset as the primary key query result; determine the second row offset range within the query range according to the row offset of the data in the primary key query result.
[0019] In one embodiment of the present specification, the data within the first local data and the second local data are stored in the form of data blocks, and each database stores partial data; the real-time query module is used to query the data within the query range in the at least one first local data and the at least one second local data according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sort the queried data as the data to be queried, and is used to: query the data blocks to which the data within the query range belongs in the at least one first local data and the at least one second local data according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sort the queried data blocks as the data to be queried.
[0020] In one embodiment of the present specification, the data blocks include macroblocks and microblocks.
[0021] In one embodiment of the present specification, the target data includes: data generated in the memory according to a data definition language DDL.
[0022] In one embodiment of the present specification, different column groups of the target data correspond to the same first partial data; or different column groups of the target data correspond to different first partial data.
[0023] According to a third aspect of one or more embodiments of this specification, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor implements the method described in the first aspect by running the executable instructions.
[0024] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0025] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:
[0026] The data processing method provided in the embodiments of this specification can, in the process of writing target data in memory to disk, respond to receiving a query instruction for the target data, query data in the at least one first local data and the at least one second local data according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sort the queried data as the data query result, wherein the first local data includes the portion of the target data stored on disk, the second local data includes the portion of the target data stored in memory, the target data is in columnar form, and each column group of the target data corresponds to one second local data. In other words, the method can provide real-time query services to the outside world during the process of writing columnar data from memory to disk; because each column group of columnar data is stored as an independent second local data in memory, the method can sort the queried data after performing data query on multiple local data in disk and memory, thereby ensuring the accuracy of the data query result. Furthermore, the method can provide query services for related data during the process of defining a data table using DDL, thereby ensuring the business processing response speed of the database and avoiding problems such as business request timeouts. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] FIG1 is a flowchart of a data processing method provided by an exemplary embodiment.
[0028] FIG2 is a schematic diagram of intermediate index layer information provided by an exemplary embodiment.
[0029] FIG3 is a flow chart of a data processing method provided by an exemplary embodiment.
[0030] FIG4 is a schematic diagram of a multi-way merge query method provided by an exemplary embodiment.
[0031] FIG. 5 is a schematic diagram of intermediate index layer information of first partial data and second partial data of a rank group C2 provided by an exemplary embodiment.
[0032] FIG. 6 is a schematic diagram of intermediate index layer information of first partial data and second partial data of a rank group C3 provided by an exemplary embodiment.
[0033] FIG7 is a schematic structural diagram of a device provided by an exemplary embodiment.
[0034] FIG8 is a block diagram of a data processing device provided by yet another exemplary embodiment. DETAILED DESCRIPTION
[0035] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.
[0036] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0037] With the rapid development of the internet and information technology, data generation is exploding, placing increasing demands on databases and their management. Data processing requires the use of DML (Data Manipulation Language) to manipulate data, such as adding, deleting, querying, and modifying it. Data processing also requires the use of DDL (Data Definition Language) to restructure data, such as creating new tables, deleting columns, and changing column types.
[0038] In the related art, when using DDL to reorganize data, the relevant data cannot provide query services to the outside world, resulting in the database's business processing being affected and causing business problems such as business request timeouts. For example, when using DDL to perform some data reorganization operations (such as deleting columns, changing column types, etc.), the reorganized data needs to be written to a hidden table in the memory and made available on disk (i.e., dumped to disk) before it can provide read and write services to the outside world. When the amount of data is large, the above process will take a long time, resulting in the reorganized data being unable to provide read and write services within a certain period of time. Especially when using DDL to reorganize column-stored data, since the reorganized data in the column store may have multiple column groups, each column group independently executes the above process (i.e., the data of each column group is written to a different hidden table and written to disk independently, i.e., the progress of writing to disk cannot be guaranteed to be synchronized). Therefore, the reorganized data is more dispersed in the above process, making it more difficult to provide read and write services to the outside world.
[0039] Based on this, on the first aspect, at least one embodiment of this specification provides a data processing method, which can provide real-time query services and other read and write services during the process of dumping data in memory to disk (i.e., the disk placement process). That is, certain data, especially column-stored data containing multiple column groups, is scattered in memory and disk during the disk placement process. This method can provide real-time read and write services for data that is in a scattered state during the disk placement process, especially column-stored data containing multiple column groups.
[0040] Please refer to FIG1 , which exemplarily shows the flow of the data processing method, including step S101 .
[0041] In step S101, in the process of writing the target data in the memory to the disk, in response to receiving a query instruction for the target data, according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, data is queried in the at least one first local data and the at least one second local data respectively, and the queried data is sorted as the data query result.
[0042] Exemplarily, the target data may include (reorganized) data generated in the memory according to the data definition language (DDL). For example, the target data may be reorganized data generated by operations such as creating a new table, deleting a column, or changing a column type according to the data definition language (DDL). The reorganized data generated according to the data definition language (DDL) must first be written to the memory and then written from the memory to the disk (i.e., flushed to disk). This step occurs during the flushing of the reorganized data to disk.
[0043] As another example, the target data can be in the form of column storage or row storage. Among them, the row storage form refers to the data storage in rows as the basic unit, and each row contains the values of all the fields in the table; the column storage form refers to the data storage in columns as the basic unit, and the data of the same column are stored together to facilitate data aggregation operations and to compress the data. Since the reorganized data stored in the column may have multiple column groups, it leads to a higher degree of dispersion during the disk transfer process (the relevant reasons have been described in detail in the previous article and will not be repeated here); therefore, the following content of the method will take column storage data as an example to introduce the process of the method, so as to improve the adaptability of the method, but this is not a limitation on the form of the target data used by the method.
[0044] The first local data includes part of the target data stored on disk (for example, it can be called SSTable).
[0045] The second local data includes part of the target data stored in the memory (for example, it may be called DDLKV).
[0046] It should be understood that if the target data is in the form of column storage, the target data includes multiple column groups (a column group is a ancestral form containing one or more columns of data in the column storage), that is, when the target data is generated in the memory, a second local data DDLKV containing multiple column groups can be generated. Preferably, each column group corresponds to a second local data DDLKV. Furthermore, in the process of writing the target data in the memory to the disk, the target data will be dispersed into at least one first local data and at least one column group of second local data. The number of first local data is determined by the form of the partial data written to the disk by each column group. For example, if the partial data written to the disk by each column group are all stored in a first local data SSTable, the number of first local data is 1, that is, different column groups of the target data correspond to the same first local data SSTable; for another example, if the partial data written to the disk by each column group are respectively stored in different first local data SSTables, the number of first local data is the same as the column group, that is, different column groups of the target data correspond to different first local data; preferably, the partial data written to the disk by each column group are all stored in a first local data SSTable.
[0047] The number of second partial data DDLKVs is determined by the progress of data flushing of each column group. For example, if all the data of a column group is flushed to disk, the second partial data DDLKV of this column group does not exist in the memory. For another example, if not all the data of a column group is flushed to disk, the second partial data DDLKV of this column group exists in the memory.
[0048] Among them, the second local data SSTalbe persisted on the disk consists of its own metadata information and a series of data macroblocks, and each data macroblock can be further divided into multiple microblocks; the data macroblock contains multiple rows of sorted data, persisted on the disk, and has a fixed size of 2M. It is the basic component unit of SSTable. The microblock is the basic component unit of the data macroblock, containing multiple rows of sorted data, persisted on the disk, and has a variable size, usually a few KB. It is the smallest unit for reading SSTable data from the disk. The first local data stored in the memory can also be stored in the memory as the second local data, that is, it consists of its own metadata and a series of data macroblocks, and each data macroblock can be further divided into multiple microblocks. The macroblocks and microblocks within the first and second local data can be organized through the intermediate index layer information to accelerate queries, that is, the intermediate index layer information of the first and second local data can be used to represent its internal data organization form, that is, to characterize the distribution information of the first local data on different macroblocks and microblocks within it. Please refer to FIG2, which exemplarily shows the intermediate index layer information of a first local data. As can be seen from the figure, the intermediate index layer information is in a tree structure, including at least one layer of index information and a bottom macroblock, and the macroblock has microblocks; the index information is divided layer by layer, that is, each layer of index information includes multiple sub-local data obtained by dividing the first local data, and the number of sub-local data contained in the lower index information is greater than the number of sub-local data contained in the upper index information. A sub-local data in the upper index information is divided into at least one sub-local data in the lower index information, and each sub-local data in the bottom index information corresponds to a data macroblock. The sub-local data can be represented by a key value range and a row offset range. For example, the sub-local data in FIG2 is represented by the end value (endkey) of the key value range and the end value of the row offset range. The row offset in the sub-local data is the absolute offset of the data in the first local data. A macroblock is composed of multiple microblocks. A microblock can be represented by a key value range and a row offset range. For example, the microblock in Figure 2 is represented by the end value of the key value range and the end value of the row offset range, but the row offset within the microblock is the relative offset of the microblock within the macroblock.
[0049] The query instruction may include a query range and a query condition. Exemplarily, this step may be performed in the manner shown in FIG. 3 , including sub-steps S1011 and S1012 .
[0050] In sub-step S1011, based on the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, data within the query range are queried in the at least one first local data and the at least one second local data respectively, and the queried data are sorted as data to be queried.
[0051] The query range may include a column query range of at least one column group, that is, the query range may be represented by the column query range of at least one column group, such as a range of values of a column in a column group. This sub-step may be performed as follows: First, for each column query range in the query range, a first row offset range corresponding to the column query range of the column group is determined based on the intermediate index layer information of the first partial data and the intermediate index layer information of the second partial data of the column group, as well as the column query range of the column group.
[0052] That is to say, the first local data SSTable and the second local data DDLKV involved in the column group can use the intermediate index layer information to simulate the intermediate index iterator to respectively execute data queries, and sort the queried data according to the row offset as the column query result of the column group, and then filter the obtained column query result according to the column query range of the column group to obtain the first row offset range corresponding to the column query range of the column group. Please refer to Figure 4. When the first local data SSTable and the second local data DDLKV involved in the column group execute data queries respectively, a multi-way merge method can be adopted to simultaneously complete data query and data sorting; for example, when the first local data SSTable or the second local data DDLKV queries a certain macroblock, if the row offset of the macroblock shows that it is the macroblock with the smallest row offset among the remaining data to be queried, then the macroblock is output to be added to the query result. First, the macroblock is not output temporarily until the macroblock is the macroblock with the smallest row offset among the remaining data to be queried. Then the macroblock is output to be added to the output result.
[0053] For example, the column query result of the column group can be determined in the following manner: based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, all the data of the column group are queried in the first local data and the second local data of the column group respectively, and the queried data are sorted according to the row offset as the column query result of the column group. For another example, the column query result of the column group can be determined in the following manner: based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, the data within the second row offset range of the query range is queried in the first local data and the second local data of the column group respectively, and the queried data is sorted according to the row offset as the column query result of the column group; wherein, the second row offset range involved in this example can be the row offset range included in the query range, or the offset range determined according to the primary key range included in the query range in the following manner: first, based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the primary key column (that is, the column where the primary key is located), the data within the primary key range is queried in the first local data and the second local data of the primary key column respectively, and the queried data is sorted according to the row offset as the primary key query result; based on the row offset of the data in the primary key query result, the second row offset range within the query range is determined (that is, the union of the offsets of each row of data in the primary key query result is determined as the second row offset range).
[0054] For example, the first row offset range corresponding to the column query range of the column group can be determined in the following manner: the union of the offsets of each row of data within the column query range that meets the column query range (the range is represented by the values within the columns of the column group) is determined as the first row offset range.
[0055] Next, the second row offset range is determined based on the first row offset range corresponding to each column query range in the query range. For example, based on the relationship between each column query range in the query range, the intersection and union of the first row offset range corresponding to each column query range are taken to obtain the second row offset range. If the query range contains three column query ranges, and the relationship between the three column query ranges is an AND relationship, then the intersection of the first row offset ranges corresponding to the three column query ranges can be determined as the second row offset range; if the query range contains three column query ranges, and the relationship between the three column query ranges is an OR relationship, then the union of the first row offset ranges corresponding to the three class query ranges can be determined as the second row offset range.
[0056] Finally, based on the second row offset range, as well as the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, the data within the second row offset range is queried in the at least one first local data and the at least one second local data, respectively, and the queried data is sorted and used as the data to be queried. In other words, the second row offset range is used to perform a multi-way merged cross-local data query on at least one first local data SSTable and at least one second local data DDLKV of the target data to achieve query-while-sorting, and ultimately output the data to be queried sorted according to the row offset.
[0057] It should be understood that the data in the first local data and the second local data are stored in the form of data blocks, which include macroblocks and microblocks. The details of macroblocks and microblocks have been described in detail in the previous text and will not be repeated here. Based on this, this sub-step can query the data blocks to which the data within the query range belongs in the at least one first local data and the at least one second local data according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sort the queried data blocks as the data to be queried, that is, at least one first local data and at least one second local data spit out at least one macroblock sorted by row offset as the query data to be queried.
[0058] In sub-step S1012, the data to be queried is queried according to the query condition to determine the data query result.
[0059] The query condition may be a query condition for at least one column. If the query condition is empty, the query data may be directly determined as the data query result.
[0060] Next, the data processing method obtained by combining the above multiple embodiments is described in detail with reference to a specific example.
[0061] Suppose there is a table t1 (c1 int primary key, c2 int, c3 int) with two column groups, cg_c2(c2) and cg_c3(c3). This table has 200 rows of data, where the data in c1 is (1...200), the data in c2 is (1001...1200), and the data in c3 is (2001...2200). The degree of parallelism during DDL execution is 2. For the row-based column group, the first thread is responsible for the data with the primary key (1...100), and the second thread is responsible for the data with the primary key (101...200). At a certain moment in the process of writing the table from memory to disk, for Column Group cg_c2, when the first thread executes to c2 is 1030, and the second thread executes to 1130, the dump generates the first first local data SSTable1, and the remaining data is stored in the second local data DDLKV1. The layout of SSTable1 and DDLKV1 can refer to Figure 5 (since the key value and row offset of each row in t1 are the same, the index information in the figure only contains the row offset); for Column Group cg_c3, when the first thread executes to c3 is 2040, and the second thread executes to 2140, the dump generates the first first local data SSTable2, and the remaining data is stored in the second local data DDLKV2. The layout of SSTable2 and DDLKV2 can refer to Figure 6 (since the key value and row offset of each row in t1 are the same, the index information in the figure only contains the row offset).
[0062] At this time, a query instruction select * from t1 where c2>1035and c3<2120and c1>=0and c1<150 is received for the table t1; the query can be completed across SSTable1, DDLKV1, SSTable2, and DDLKV2.
[0063] First: according to the primary key range [0,150) in the query instruction, and the intermediate index layer information of SSTable1 and the intermediate index layer information of DDLKV1 shown in Figure 5 or the intermediate index layer information of SSTable2 and the intermediate index layer information of DDLKV2 shown in Figure 6, the intermediate index iterator can be used to iterate out [25,0,25], [50,26,50], [75,51,75], [100,76,100], [125,101,125], [150,126,150] in sequence (Note: the above brackets are the key value end value, row offset starting value, and row offset ending value respectively). Therefore, the second row offset range row_offset in the query instruction can be obtained as: start_row_offset is 0, and end_row_offset is 150.
[0064] Next, based on the second row offset range row_offset, Column Group C2 and Column Group C3 are queried respectively. For the query of Column Group C2, according to the second row offset range row_offset, and the intermediate index layer information of SSTable1 and the intermediate index layer information of DDLKV1 shown in Figure 5, the intermediate index iterator is used to iterate out [0,15], [16,30], [31,45], [46,100], [101,115], [116,130], [131,145], [146,150] in sequence (Note: the values in the square brackets are the starting value and the ending value of the row offset respectively); in the above iteration process, apply_filter can be used according to a certain batch size. Since the column query range of Column Group C2 is >1035, in the result map generated by Column Group C2, the front [0,35] is 0 and the back [36,150] is 1, and the column query result of Column Group C2 is [36,150]. For the query of Column Group C3, according to the second row offset range row_offset, and the intermediate index layer information of SSTable2 and the intermediate index layer information of DDLKV2 shown in Figure 6, the intermediate index iterator is used to iterate out [0,15], [16,40], [41,55], [56,100], [101,115], [116,140], [141,150] in sequence (Note: the values in the square brackets are the starting value and the ending value of the row offset respectively); in the above iteration process, apply_filter can be applied according to a certain batch size. Since the column query range of Column Group C3 is <2020, in the result map generated by Column Group C3, the front [0,119] is 1 and the back [120,150] is 0, that is, the column query result of Column Group C3 is [0,119].
[0065] Next, since the column query ranges of C2 and C3 are in the form of an AND, the intersection of the column query range [36,150] of Column Group C2 and the column query result [0,119] of Column Group C3 is taken to obtain the first row offset range [36,119].
[0066] Finally, based on the first row offset range [36,119], the intermediate index rows are iterated in Column Group C2 and Column Group C3 respectively to determine the microblocks involved in the query data. Then, based on the microblocks involved in the query data and the query condition result bitmap, it is determined whether a row of the microblock should be projected into the data query result.
[0067] The data processing method provided in the embodiments of this specification can, in the process of writing target data in memory to disk, respond to receiving a query instruction for the target data, query data in the at least one first local data and the at least one second local data according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sort the queried data as the data query result, wherein the first local data includes the portion of the target data stored on disk, the second local data includes the portion of the target data stored in memory, the target data is in columnar form, and each column group of the target data corresponds to one second local data. In other words, the method can provide real-time query services to the outside world during the process of writing columnar data from memory to disk; because each column group of columnar data is stored as an independent second local data in memory, the method can sort the queried data after performing data query on multiple local data in disk and memory, thereby ensuring the accuracy of the data query result. Furthermore, the method can provide query services for related data during the process of defining a data table using DDL, thereby ensuring the business processing response speed of the database and avoiding problems such as business request timeouts.
[0068] The data processing method provided in the embodiments of this specification has different progress in the process of reorganizing data in column storage form from memory to disk. In order to avoid playback jams and provide external query services as soon as possible, a fusion query solution can be proposed based on the middle-layer index information. By fusing and merging the queried data in the middle-layer index iterator, the required data can be iterated accurately and efficiently.
[0069] FIG7 is a schematic structural diagram of a device provided by an exemplary embodiment. Referring to FIG7 , at the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710, and may also include hardware required for other tasks. One or more embodiments of this specification may be implemented based on software, such as the processor 702 reading the corresponding computer program from the non-volatile memory 710 into the memory 708 and then running it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but may also be hardware or logic devices.
[0070] Referring to FIG8 , a data processing device can be applied to the device shown in FIG7 to implement the technical solution of this specification. The device includes: a real-time query module 801 for, in the process of writing target data in memory to disk, responding to receiving a query instruction for the target data, querying data in the at least one first local data and the at least one second local data according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sorting the queried data as a data query result; wherein the first local data includes the portion of the target data stored on disk, the second local data includes the portion of the target data stored in memory, the target data is stored in column format, and each column group of the target data corresponds to one second local data.
[0071] In one embodiment of the present specification, the query instruction includes a query range and a query condition; the real-time query module is used to: query the data within the query range in the at least one first local data and the at least one second local data according to the intermediate index layer information of the at least one first local data and the intermediate index layer information of the at least one second local data, and sort the queried data as the data to be queried; query the data to be queried according to the query condition to determine the data query result.
[0072] In one embodiment of the present specification, the query range includes a column query range of at least one column group; the real-time query module is used to query the data within the query range in the at least one first local data and the at least one second local data according to the query range, the intermediate index layer information of at least one first local data, and the intermediate index layer information of at least one second local data, and sort the queried data as the data to be queried, and is used to: for the column group to which each column query range in the query range belongs, determine the first row offset range corresponding to the column query range of the column group according to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, and the column query range of the column group; determine the second row offset range according to the first row offset range corresponding to each column query range in the query range; query the data within the second row offset range in the at least one first local data and the at least one second local data according to the second row offset range, the intermediate index layer information of at least one first local data, and the intermediate index layer information of at least one second local data, and sort the queried data as the data to be queried.
[0073] In one embodiment of the present specification, the real-time query module is used to determine the first row offset range corresponding to the column query range of the column group based on the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, and the column query range of the column group. It is used to: query the first local data and the second local data of the column group for data within the third row offset range of the query range according to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the column group, respectively, and sort the queried data according to the row offset as the column query result of the column group; filter the column query result of the column group according to the column query range of the column group to obtain the first row offset range corresponding to the column query range of the column group.
[0074] In one embodiment of the present specification, the query range includes a primary key range; the device also includes a primary key module, which is used to: query the data within the primary key range in the first local data and the second local data of the primary key column according to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the primary key column, and sort the queried data according to the row offset as the primary key query result; determine the second row offset range within the query range according to the row offset of the data in the primary key query result.
[0075] In one embodiment of the present specification, the data within the first local data and the second local data are stored in the form of data blocks, and each database stores partial data; the real-time query module is used to query the data within the query range in the at least one first local data and the at least one second local data according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sort the queried data as the data to be queried, and is used to: query the data blocks to which the data within the query range belongs in the at least one first local data and the at least one second local data according to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sort the queried data blocks as the data to be queried.
[0076] In one embodiment of the present specification, the data blocks include macroblocks and microblocks.
[0077] In one embodiment of the present specification, the target data includes: data generated in the memory according to a data definition language DDL.
[0078] In one embodiment of the present specification, different column groups of the target data correspond to the same first partial data; or different column groups of the target data correspond to different first partial data.
[0079] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.
[0080] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0081] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0082] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0083] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0084] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0085] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "an," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0086] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0087] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when..." or "when..." or "in response to determining."
[0088] The above description is one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included in the scope of protection of the claims of this application.
Claims
1. A data processing method, comprising: In the process of writing the target data in the memory to the disk, in response to receiving a query instruction for the target data, according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively query data in the at least one first local data and the at least one second local data, and sort the queried data as the data query result; The first local data includes part of the target data stored in the disk, the second local data includes part of the target data stored in the memory, the target data is in column storage, and each column group of the target data corresponds to a second local data.
2. The data processing method according to claim 1, wherein: The query instruction includes a query range and a query condition; The step of searching for data in the at least one first local data and the at least one second local data respectively according to the query instruction and the intermediate index layer information of the at least one first local data and the intermediate index layer information of the at least one second local data, and sorting the queried data as the data query result includes: According to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively query the at least one first local data and the at least one second local data for data within the query range, and sort the queried data as data to be queried; The data to be queried is queried according to the query condition to determine the data query result.
3. The data processing method according to claim 2, wherein: The query range includes a column query range of at least one column group; The method of searching for data within the query range in the at least one first local data and the at least one second local data respectively according to the query range and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, and sorting the queried data as data to be queried includes: For the column group to which each column query range in the query range belongs, determining a first row offset range corresponding to the column query range of the column group according to the intermediate index layer information of the first local data of the column group and the intermediate index layer information of the second local data, and the column query range of the column group; Determine the second row offset range according to the first row offset range corresponding to each column query range in the query range; According to the second row offset range, and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively, in the at least one first local data and the at least one The second local data is searched for data within the second row offset range, and the searched data is sorted as the data to be searched.
4. The data processing method according to claim 3, wherein: The determining, according to the intermediate index layer information of the first partial data of the column group and the intermediate index layer information of the second partial data, and the column query range of the column group, a first row offset range corresponding to the column query range of the column group includes: According to the intermediate index layer information of the first partial data and the intermediate index layer information of the second partial data of the column group, respectively query the first partial data and the second partial data of the column group for data within the third row offset range of the query range, and sort the queried data according to the row offset as the column query result of the column group; The column query results of the column group are screened according to the column query range of the column group to obtain a first row offset range corresponding to the column query range of the column group.
5. The data processing method according to claim 4, wherein: The query range includes the primary key range; The method further comprises: According to the intermediate index layer information of the first local data and the intermediate index layer information of the second local data of the primary key column, respectively query the data within the primary key range in the first local data and the second local data of the primary key column, and sort the queried data according to the row offset as the primary key query result; According to the row offset of the data in the primary key query result, a second row offset range within the query range is determined.
6. The data processing method according to claim 2, wherein: The data in the first local data and the second local data are stored in the form of data blocks, and each database stores part of the data; The method of searching for data within the query range in the at least one first local data and the at least one second local data respectively according to the intermediate index layer information of the at least one first local data and the at least one second local data, and sorting the queried data as the data to be queried includes: According to the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, the data blocks to which the data within the query range belongs are queried in the at least one first local data and the at least one second local data respectively, and the queried data blocks are sorted as the data to be queried.
7. The data processing method according to claim 6, wherein: The data blocks include macroblocks and microblocks.
8. The data processing method according to claim 1, wherein: The target data includes: The data is generated in the memory according to a data definition language DDL.
9. The data processing method according to claim 1, wherein: Different column groups of the target data correspond to the same first partial data; or different column groups of the target data correspond to different first partial data.
10. A data processing device, comprising: The real-time query module is used to respond to the received query in the process of writing the target data in the memory to the disk. The query instruction of the target data, according to the query instruction and the intermediate index layer information of at least one first local data and the intermediate index layer information of at least one second local data, respectively querying data in the at least one first local data and the at least one second local data, and sorting the queried data as the data query result; The first local data includes part of the target data stored in the disk, the second local data includes part of the target data stored in the memory, the target data is in column storage, and each column group of the target data corresponds to a second local data.
11. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 9 by running the executable instructions.
12. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Real-time index creating and real-time searching method and device
CN103294731A
Data searching method ad device
CN108846121A
Data processing method and device for distributed database
CN115658811A
Real-time object storage query system based on search engine and relational database
CN116204556A
Data processing method and device, electronic equipment and storage medium
CN117688033A