Data storage structure processing method, processing device and data storage system

By creating a data parent table and splitting it into row storage and column storage child tables in the HTAP scenario, the problem of poor performance in inserting data and querying data in the HTAP scenario is solved, and efficient data processing and resource utilization are achieved.

CN114238318BActive Publication Date: 2025-05-06中国邮政储蓄银行股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111471251.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-05-06
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The performance of inserting data and querying data in HTAP scenarios is poor. The existing technology is processed through table sub-table, but it leads to high consumption of IO resources, low retrieval efficiency, inability to create updateable views, and performance is degraded when processing a large amount of concurrent data.

Method used

By creating a data parent table, including index information for mapping row storage subtables and column storage subtables, the row storage subtables are used to store data of the first target time period, the column storage subtables are used to store previous data, and the row storage subtables are used to store row storage secondary subtables and column storage secondary subtables according to the split period, which are used to store data for different time periods.

Benefits of technology

It realizes that there is both a row storage structure and a column storage structure in a table, which improves query efficiency, reduces hardware resource occupation and I/O resource consumption, and improves the efficiency of inserting and updating data, avoiding the difficulty and high coupling of the writing logic process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114238318B_ABST
    Figure CN114238318B_ABST
Patent Text Reader

Abstract

The present application provides a data storage structure processing method, a processing device and a data storage system, the method comprising: creating a data parent table, the data parent table comprising index information, the index information being used to characterize a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, the row storage sub-table being used to store data in a first target time period, and the column storage sub-table being used to store data before the first target time period; determining whether to split the row storage sub-table according to a splitting cycle; and in the case of determining to split the row storage sub-table, splitting the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table being used to store data in a second target time period, and the column storage secondary sub-table being used to store data in a first target time period, the second target time period being a time period after the first target time period, thereby solving the problem of poor performance of inserting and querying data in an HTAP scenario in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data storage, and in particular, to a data storage structure processing method, a processing device, a computer-readable storage medium, a processor, and a data storage system. Background Art

[0002] Distributed databases are high-performance, highly reliable, and easily scalable databases that have emerged in recent years due to the growth of data, increased business complexity, and increased user traffic. Distributed databases shard data according to certain rules and use a method of breaking the whole into parts to optimize the processing of performance issues caused by excessive data volume and concurrency. It has a certain number of autonomous processing units that are interconnected through the network and work together to process their assigned tasks, and finally submit them to the user interface in a unified manner.

[0003] The business scenarios processed by distributed databases are divided into three categories: one is the OLTP (On-Line Transaction Processing, OLTP for short) scenario for transactional businesses, such as Taobao's shopping system, bank deposit and remittance systems, 12306's ticket purchasing system, etc.; one is the OLAP (On-Line Analytical Processing, OLAP for short) scenario for analytical businesses, such as various industry data reports and business intelligence reports published by the National Bureau of Statistics; and one is the HTAP (Hybrid Transaction Analytical Processing, HTAP for short) scenario for hybrid businesses, such as today's headlines news recommendations, bank mobile APP transaction query and income and expenditure analysis, etc.

[0004] For HTAP scenarios, traditional relational databases usually adopt solutions similar to split tables. For example, in Oracle, some tables are created as row-stored tables, and other tables are created as column-stored tables. When querying, a result set of multiple tables is created through views, and a unified query object is provided to the outside world. When inserting or updating, the application needs to be modified to write the data into the corresponding table.

[0005] Although the use of sharding can also solve the functional problems of HTAP scenarios, it also brings various other negative effects. For example, when merging tables with large data volumes, a large amount of IO resources and memory will be consumed; when using the merged table for conditional filtering, the existing indexes on the table are usually not available, resulting in very low retrieval efficiency; merged views cannot create updateable views; and since the judgment logic needs to be called every time a piece of data is added or modified, for application systems that handle large amounts of concurrency and data, the performance may be reduced exponentially.

[0006] Therefore, there is an urgent need for a method that can optimize the performance of inserting and querying data in HTAP scenarios.

[0007] The above information disclosed in the background technology section is only used to enhance the understanding of the background technology of the technology described in this article. Therefore, the background technology may contain certain information that does not form the prior art known in this country for those skilled in the art. Summary of the invention

[0008] The main purpose of the present application is to provide a data storage structure processing method, a processing device, a computer-readable storage medium, a processor and a data storage system to solve the problem of poor performance of inserting and querying data in HTAP scenarios in the prior art.

[0009] According to one aspect of an embodiment of the present invention, a method for processing a data storage structure is provided, including: creating a data parent table, the data parent table including index information, the index information being used to characterize a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, the row storage sub-table being used to store data in a first target time period, and the column storage sub-table being used to store data before the first target time period; determining whether to split the row storage sub-table according to a splitting cycle; and if it is determined that the row storage sub-table is to be split, splitting the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table being used to store data in a second target time period, and the column storage secondary sub-table being used to store data in the first target time period, the second target time period being a time period after the first target time period.

[0010] Optionally, in the case of determining to split the row storage sub-table, after splitting the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table, the method further includes: determining whether to split the row storage secondary sub-table according to the splitting cycle; in the case of determining to split the row storage secondary sub-table, splitting the row storage secondary sub-table into a row storage tertiary sub-table and a column storage tertiary sub-table, the row storage tertiary sub-table being used to store data of a third target time period, and the column storage tertiary sub-table being used to store data of the second target time period, the third target time period being a time period after the second target time period.

[0011] Optionally, the row storage secondary subtable is used to process data in an OLTP scenario.

[0012] Optionally, the column storage secondary subtable is used to process data in an OLAP scenario.

[0013] Optionally, the data parent table also includes multiple field information, the row storage sub-table and the column storage sub-table inherit the multiple field information of the data parent table, and the row storage secondary sub-table and the column storage secondary sub-table inherit the multiple field information of the row storage sub-table.

[0014] Optionally, the method also includes: controlling the data parent table to receive target request information, wherein the target request information is request information for inserting, deleting, changing or querying data; according to the index information, controlling the data parent table to send the target request information to the corresponding storage sub-table, and controlling the data parent table to send the storage sub-table's response information to the target request information to the application.

[0015] According to another aspect of an embodiment of the present invention, a processing device for a data storage structure is also provided, including: a creation unit, used to create a data parent table, the data parent table including index information, the index information being used to characterize a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, the row storage sub-table being used to store data in a first target time period, and the column storage sub-table being used to store data before the first target time period; a first determination unit, used to determine whether to split the row storage sub-table according to a splitting cycle; and a first splitting unit, used to split the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table when it is determined that the row storage sub-table is to be split, the row storage secondary sub-table being used to store data in a second target time period, and the column storage secondary sub-table being used to store data in the first target time period, the second target time period being a time period after the first target time period.

[0016] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein the program executes any one of the methods described.

[0017] According to yet another aspect of the embodiments of the present invention, a processor is provided, wherein the processor is used to run a program, wherein any one of the methods is executed when the program is run.

[0018] According to one aspect of an embodiment of the present invention, there is also provided a data storage system, comprising: one or more processors, a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of the described methods.

[0019] In an embodiment of the present invention, in a method for processing a data storage structure, first, a data parent table is created, the data parent table including index information for characterizing a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, wherein the row storage sub-table is used to store data in a first target time period, and the column storage sub-table is used to store data before the first target time period; secondly, according to a splitting cycle, it is determined whether to split the row storage sub-table; finally, in the case of determining to split the row storage sub-table, the row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table is used to store data in a second target time period, and the column storage secondary sub-table is used to store data in the first target time period, and the second target time period is a time period after the first target time period. In this solution, the row storage subtable is split into a row storage secondary subtable and a column storage secondary subtable. The row storage secondary subtable is used to store data of the second target time period, that is, the row storage secondary subtable is used to store data close to the current time, and the column storage secondary subtable is used to store data far from the current time. This ensures that the obtained row storage secondary subtable is used to store the latest data, and the added column storage secondary subtable will not affect the hardware resources of the system. In addition, this solution implements the existence of both a row storage structure and a column storage structure in one table. Compared with the prior art that uses a multi-table merge method to process query data in the HTAP scenario, this solution performs queries in one table without the need to merge multiple tables. This not only ensures that less hardware resources are occupied and less I / O resources are consumed, but also ensures that the query efficiency is high by querying based on index information. In addition, when inserting or updating data, this solution does not need to write a logical process through an application program, which ensures that the efficiency of inserting and updating data is high, and also avoids the problem of high difficulty and high coupling in writing logical processes, thereby solving the problem of poor performance of inserting and querying data in the HTAP scenario in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings constituting part of the present application are used to provide a further understanding of the present application. The exemplary embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0021] Figure 1 A schematic diagram showing a method for processing a data storage structure according to an embodiment of the present application is shown;

[0022] Figure 2 A schematic diagram showing a processing device for a data storage structure according to an embodiment of the present application is shown;

[0023] Figure 3A storage relationship diagram of a data parent table according to an embodiment of the present application is shown;

[0024] Figure 4 A schematic diagram of creating a data parent table according to an embodiment of the present application is shown;

[0025] Figure 5 A schematic diagram showing automatic splitting of a row storage sub-table according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0026] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0027] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] For the convenience of description, some nouns or terms involved in the embodiments of the present application are explained below:

[0030] OLAP: On-Line Analytical Processing, which enables analysts to observe information from all aspects quickly, consistently, and interactively to achieve a deep understanding of the data;

[0031] OLTP: On-Line Transaction Processing, is a transaction-oriented process. Its basic feature is that the user data received by the front desk can be immediately transmitted to the computing center for processing, and the processing results can be given in a very short time. It is one of the ways to quickly respond to user operations.

[0032] HTAP: Hybrid Transaction Analytical Processing, is the abbreviation of online transaction processing and online analytical processing;

[0033] Database: A collection of data organized according to a predetermined structure;

[0034] Distributed database: Logically it is a unified whole, but physically it is a database cluster stored on different physical nodes.

[0035] As mentioned in the background technology, the performance of inserting and querying data in the HTAP scenario in the prior art is poor. In order to solve the above problems, in a typical implementation of the present application, a data storage structure processing method, a processing device, a computer-readable storage medium, a processor and a data storage system are provided.

[0036] According to an embodiment of the present application, a method for processing a data storage structure is provided.

[0037] Figure 1 1 is a flowchart of a method for processing a data storage structure according to an embodiment of the present application. Figure 1 As shown, the method comprises the following steps:

[0038] Step S101, creating a data parent table, the data parent table including index information, the index information being used to characterize a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, the row storage sub-table being used to store data in a first target time period, and the column storage sub-table being used to store data before the first target time period;

[0039] Step S102, determining whether to split the row storage sub-table according to the splitting cycle;

[0040] Step S103, when it is determined that the row storage sub-table is to be split, the row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table is used to store data of the second target time period, and the column storage secondary sub-table is used to store data of the first target time period, and the second target time period is a time period after the first target time period.

[0041] In the processing method of the above data storage structure, firstly, a data parent table is created, and the above data parent table includes index information for characterizing the mapping relationship between the above data parent table and the row storage sub-table and the column storage sub-table, wherein the above row storage sub-table is used to store data in the first target time period, and the above column storage sub-table is used to store data before the above first target time period; secondly, according to the splitting cycle, it is determined whether to split the above row storage sub-table; finally, in the case of determining to split the above row storage sub-table, the above row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the above row storage secondary sub-table is used to store data in the second target time period, and the above column storage secondary sub-table is used to store data in the first target time period, and the second target time period is a time period after the above first target time period. In this solution, the row storage subtable is split into a row storage secondary subtable and a column storage secondary subtable. The row storage secondary subtable is used to store data of the second target time period, that is, the row storage secondary subtable is used to store data close to the current time, and the column storage secondary subtable is used to store data far from the current time. This ensures that the obtained row storage secondary subtable is used to store the latest data, and the added column storage secondary subtable will not affect the hardware resources of the system. In addition, this solution implements the existence of both a row storage structure and a column storage structure in one table. Compared with the prior art that uses a multi-table merge method to process query data in the HTAP scenario, this solution performs queries in one table without the need to merge multiple tables. This not only ensures that less hardware resources are occupied and less I / O resources are consumed, but also ensures that the query efficiency is high by querying based on index information. In addition, when inserting or updating data, this solution does not need to write a logical process through an application program, which ensures that the efficiency of inserting and updating data is high, and also avoids the problem of high difficulty and high coupling in writing logical processes, thereby solving the problem of poor performance of inserting and querying data in the HTAP scenario in the prior art.

[0042] Specifically, the above-mentioned data parent table is the entrance and exit of all data. The data parent table inherits multiple child tables downward. All child tables have the table structure of the data parent table, but the storage mode is different. Some child tables use the row storage mode, and some child tables use the column storage mode. The data parent table does not store data, but only records the mapping relationship with the row storage child table and the column storage child table. The data parent table can be accessed and operated through the application. The application only needs to send the target request information to the data parent table, and access and modify the relevant data through the sharding rules of the data parent table corresponding to the child table. After creating index information on the data parent table, all child tables will also create corresponding index information. For example, based on the creation of index information based on the data parent table T1, index information can be created for its child tables T1_P1, T1_P2 to T1_Pn respectively. Only the child table has specific index data, while the data parent table T1 only has index information of the index.

[0043] In actual application, the splitting cycle may be one day, but is not limited to one day. The splitting cycle may also be determined according to actual application scenarios.

[0044] Specifically, Figure 3 As shown, when the above data parent table is created, the above data parent table includes at least one row storage sub-table and one column storage sub-table.

[0045] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0046] In one embodiment of the present application, when it is determined that the row storage sub-table is to be split, after the row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the method further includes: determining whether to split the row storage secondary sub-table according to the splitting cycle; when it is determined that the row storage secondary sub-table is to be split, splitting the row storage secondary sub-table into a row storage tertiary sub-table and a column storage tertiary sub-table, the row storage tertiary sub-table being used to store data for a third target time period, the column storage tertiary sub-table being used to store data for the second target time period, the third target time period being a time period after the second target time period. In this embodiment, whether to split the row storage level 2 sub-table is determined according to the splitting cycle. If splitting is determined, the row storage level 2 sub-table is split into a row storage level 3 sub-table and a column storage level 3 sub-table, wherein the row storage level 3 sub-table is used to store data of the third target time period, and the column storage level 3 sub-table is used to store data of the second target time period. This ensures that the row storage level 3 sub-table obtained by splitting can always be applicable to OLTP scenarios and the latest data, and the increasing number of column storage level 3 sub-tables will not affect hardware resources.

[0047] It should be noted that, as time goes by, the row-stored third-level sub-table in this application can be continuously split to obtain a row-stored fourth-level sub-table and a column-stored fourth-level sub-table. The row-stored fourth-level sub-table is generally used to store the data of the day, that is, the latest, real-time data, suitable for OLTP data, and the column-stored fourth-level sub-table is generally used to store the current data, that is, it can be understood as historical data, to adapt to OLAP scenarios.

[0048] Specifically, the storage sub-table (i.e., the row storage sub-table, column storage sub-table, and the sub-tables obtained by subsequent continuous splitting) is the physical table where all data is actually stored. The data sub-table only stores a part of the data of the entire table, which can be determined by different time periods. The data storage structure can choose row storage or column storage accordingly. The data shards that are closer to the current time point usually use row storage to process OLTP scenario business; the data shards that are farther from the current time point usually use column storage to process OLAP scenario business. When a specific operation is transmitted to the data parent table through the application, the data parent table automatically routes the operation to the corresponding storage sub-table according to the sharding rules, and performs step-by-step processing. Finally, the data parent table is summarized and returned to the application. For example, when updating a piece of data based on the data parent table T1, the data rule is used to find the T1_P1 child table where the data is located and update it, so there is no need to scan the data in T1_P2 to T1_Pn. For another example, based on the data parent table T1, an index scan is performed using a non-sharding key. At this time, index scans can be performed on T1_P1, T1_P2 to T1_Pn respectively, and the acquired data can be merged, and finally the data can be returned to the application through the data parent table. These operations greatly reduce the I / O resources occupied by the scan.

[0049] In another embodiment of the present application, the above-mentioned row storage secondary sub-table is used to process data in the OLTP scenario. In this solution, the row storage secondary sub-table is stored according to the row data as the basic logical unit, and the data of each row exists in the storage medium in a continuous storage form. Its advantage is that for random addition, deletion, modification and query operations, the row storage mode can quickly locate the data and perform the corresponding operations during the operation process. Since the amount of data in the OLTP scenario is small, but the data query and update are more frequent, the row storage secondary sub-table is used to process the data in the OLTP scenario, which not only ensures that the occupied I / O resources are less, but also ensures better data processing performance.

[0050] In another embodiment of the present application, the above-mentioned column storage secondary subtable is used to process data in an OLAP scenario. In this solution, the column storage secondary subtable is stored in a column-based logical unit, and the data of each column is stored in a storage medium in a continuous storage form. Its advantage is that for access query operations of a small number of columns, only the data of the relevant columns need to be accessed during the operation. Since data in an OLAP scenario is generally used for analysis and support of management decisions, the column storage secondary subtable is used in this solution to process data in an OLAP scenario, which not only further reduces I / O resources, but also reduces the storage space of data through compression algorithms.

[0051] In order to further ensure high efficiency of inserting and querying data, in another embodiment of the present application, the above-mentioned data parent table also includes multiple field information, the above-mentioned row storage sub-table and the above-mentioned column storage sub-table inherit the multiple field information of the above-mentioned data parent table, and the above-mentioned row storage secondary sub-table and the above-mentioned column storage secondary sub-table inherit the multiple field information of the above-mentioned row storage sub-table.

[0052] Specifically, the above-mentioned data parent table also includes multiple field information, and the above-mentioned multiple field information can be understood as the attribute information of multiple attributes in the above-mentioned data parent table. For example, the above-mentioned field information can be age, gender, date of birth, etc., but is not limited to these exemplified field information. The row storage sub-table and the column storage sub-table inherit the multiple field information of the above-mentioned data parent table, and the row storage secondary sub-table and the column storage secondary sub-table inherit the multiple field information of the above-mentioned row storage sub-table. This ensures that the storage sub-tables obtained by subsequent splitting have the same field information as the data parent table, which ensures that data can be inserted and queried more conveniently, and further ensures high efficiency in querying and inserting data.

[0053] In one embodiment of the present application, the method further includes: controlling the data parent table to receive target request information, the target request information being request information for inserting, deleting, changing or querying data; controlling the data parent table to send the target request information to the corresponding storage sub-table according to the index information, and controlling the data parent table to send the response information of the storage sub-table to the target request information to the application. In this embodiment, according to the index information, the data parent table is controlled to send the target request information to the corresponding storage sub-table, that is, the corresponding storage sub-table is scanned and searched according to the index information, without querying other storage sub-tables, thus ensuring that the efficiency of scanning and searching the storage sub-table is high, and then the data parent table is controlled to send the response information of the storage sub-table to the target request information to the application, further ensuring that the occupied I / O resources are less.

[0054] It should be noted that the above storage sub-tables may include row storage sub-tables, column storage sub-tables, row storage secondary sub-tables, column storage secondary sub-tables, row storage tertiary sub-tables, column storage tertiary sub-tables, and subsequently split storage sub-tables.

[0055] The embodiment of the present application also provides a data storage structure processing device. It should be noted that the data storage structure processing device of the embodiment of the present application can be used to execute the data storage structure processing method provided by the embodiment of the present application. The following introduces the data storage structure processing device provided by the embodiment of the present application.

[0056] Figure 2 Schematic diagram of a processing device for a data storage structure according to an embodiment of the present application. Figure 2 As shown, the device comprises:

[0057] A creation unit 10 is used to create a data parent table, the data parent table includes index information, the index information is used to represent the mapping relationship between the data parent table and the row storage sub-table and the column storage sub-table, the row storage sub-table is used to store data in the first target time period, and the column storage sub-table is used to store data before the first target time period;

[0058] A first determining unit 20 is used to determine whether to split the row storage sub-table according to the splitting cycle;

[0059] The first splitting unit 30 is used to split the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table when it is determined to split the row storage sub-table, the row storage secondary sub-table is used to store data of the second target time period, and the column storage secondary sub-table is used to store data of the first target time period, and the second target time period is a time period after the first target time period.

[0060] In the processing device of the above-mentioned data storage structure, the creation unit is used to create a data parent table, the above-mentioned data parent table includes index information, the above-mentioned index information is used to characterize the mapping relationship between the above-mentioned data parent table and the row storage sub-table and the column storage sub-table, the above-mentioned row storage sub-table is used to store data in the first target time period, and the above-mentioned column storage sub-table is used to store data before the above-mentioned first target time period; the first determination unit is used to determine whether to split the above-mentioned row storage sub-table according to the splitting cycle; the first splitting unit is used to split the above-mentioned row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table when it is determined that the above-mentioned row storage sub-table is split, the above-mentioned row storage secondary sub-table is used to store data in the second target time period, and the above-mentioned column storage secondary sub-table is used to store data in the first target time period, and the above-mentioned second target time period is a time period after the above-mentioned first target time period. In this solution, the row storage subtable is split into a row storage secondary subtable and a column storage secondary subtable. The row storage secondary subtable is used to store data of the second target time period, that is, the row storage secondary subtable is used to store data close to the current time, and the column storage secondary subtable is used to store data far from the current time. This ensures that the obtained row storage secondary subtable is used to store the latest data, and the added column storage secondary subtable will not affect the hardware resources of the system. In addition, this solution implements the existence of both a row storage structure and a column storage structure in one table. Compared with the prior art that uses a multi-table merge method to process query data in the HTAP scenario, this solution performs queries in one table without the need to merge multiple tables. This not only ensures that less hardware resources are occupied and less I / O resources are consumed, but also ensures that the query efficiency is high by querying based on index information. In addition, when inserting or updating data, this solution does not need to write a logical process through an application program, which ensures that the efficiency of inserting and updating data is high, and also avoids the problem of high difficulty and high coupling in writing logical processes, thereby solving the problem of poor performance of inserting and querying data in the HTAP scenario in the prior art.

[0061] Specifically, the above-mentioned data parent table is the entrance and exit of all data. The data parent table inherits multiple child tables downward. All child tables have the table structure of the data parent table, but the storage mode is different. Some child tables use the row storage mode, and some child tables use the column storage mode. The data parent table does not store data, but only records the mapping relationship with the row storage child table and the column storage child table. The data parent table can be accessed and operated through the application. The application only needs to send the target request information to the data parent table, and access and modify the relevant data through the sharding rules of the data parent table corresponding to the child table. After creating index information on the data parent table, all child tables will also create corresponding index information. For example, based on the creation of index information based on the data parent table T1, index information can be created for its child tables T1_P1, T1_P2 to T1_Pn respectively. Only the child table has specific index data, while the data parent table T1 only has index information of the index.

[0062] In actual application, the splitting cycle may be one day, but is not limited to one day. The splitting cycle may also be determined according to actual application scenarios.

[0063] Specifically, Figure 3 As shown, when the above data parent table is created, the above data parent table includes at least one row storage sub-table and one column storage sub-table.

[0064] In one embodiment of the present application, after splitting the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table when it is determined that the row storage sub-table is split, the device further includes a second determination unit and a second splitting unit, wherein the second determination unit is used to determine whether to split the row storage secondary sub-table according to the splitting cycle; and the second splitting unit is used to split the row storage secondary sub-table into a row storage tertiary sub-table and a column storage tertiary sub-table when it is determined that the row storage secondary sub-table is split, wherein the row storage tertiary sub-table is used to store data for a third target time period, and the column storage tertiary sub-table is used to store data for the second target time period, and the third target time period is a time period after the second target time period. In this embodiment, whether to split the row storage level 2 sub-table is determined according to the splitting cycle. If splitting is determined, the row storage level 2 sub-table is split into a row storage level 3 sub-table and a column storage level 3 sub-table, wherein the row storage level 3 sub-table is used to store data of the third target time period, and the column storage level 3 sub-table is used to store data of the second target time period. This ensures that the row storage level 3 sub-table obtained by splitting can always be applicable to OLTP scenarios and the latest data, and the increasing number of column storage level 3 sub-tables will not affect hardware resources.

[0065] It should be noted that, as time goes by, the row-stored third-level sub-table in this application will continue to split into a row-stored fourth-level sub-table and a column-stored fourth-level sub-table. The row-stored fourth-level sub-table is generally used to store the data of the day, that is, the latest, real-time data, which is suitable for OLTP data. The column-stored fourth-level sub-table is generally used to store data before the current time, which can be understood as historical data and is used to adapt to OLAP scenarios.

[0066] Specifically, the storage sub-table (i.e., the row storage sub-table, column storage sub-table, and the sub-tables obtained by subsequent continuous splitting) is the physical table where all data is actually stored. The storage sub-table only stores a part of the data of the entire table, which can be determined by different time periods. The data storage structure can choose row storage or column storage accordingly. The data shards that are closer to the current time point usually use row storage to process OLTP scenario business; the data shards that are farther from the current time point usually use column storage to process OLAP scenario business. When a specific operation is transmitted to the data parent table through the application, the data parent table automatically routes the operation to the corresponding storage sub-table according to the sharding rules, and performs step-by-step processing. Finally, the data parent table is summarized and returned to the application. For example, when updating a piece of data based on the data parent table T1, the data rule is used to find the T1_P1 child table where the data is located and update it, so there is no need to scan the data in T1_P2 to T1_Pn. For another example, based on the data parent table T1, an index scan is performed using a non-sharding key. At this time, index scans can be performed on T1_P1, T1_P2 to T1_Pn respectively, and the acquired data can be merged, and finally the data can be returned to the application through the data parent table. These operations greatly reduce the I / O resources occupied by the scan.

[0067] In another embodiment of the present application, the above-mentioned row storage secondary sub-table is used to process data in the OLTP scenario. In this solution, the row storage secondary sub-table is stored according to the row data as the basic logical unit, and the data of each row exists in the storage medium in a continuous storage form. Its advantage is that for random addition, deletion, modification and query operations, the row storage mode can quickly locate the data and perform the corresponding operations during the operation process. Since the amount of data in the OLTP scenario is small, but the data query and update are more frequent, the row storage secondary sub-table is used to process the data in the OLTP scenario, which not only ensures that the occupied I / O resources are less, but also ensures better data processing performance.

[0068] In another embodiment of the present application, the above-mentioned column storage secondary subtable is used to process data in an OLAP scenario. In this solution, the column storage secondary subtable is stored in a column-based logical unit, and the data of each column is stored in a storage medium in a continuous storage form. Its advantage is that for access query operations of a small number of columns, only the data of the relevant columns need to be accessed during the operation. Since data in an OLAP scenario is generally used for analysis and support of management decisions, the column storage secondary subtable is used in this solution to process data in an OLAP scenario, which not only further reduces I / O resources, but also reduces the storage space of data through compression algorithms.

[0069] In order to further ensure high efficiency of inserting and querying data, in another embodiment of the present application, the above-mentioned data parent table also includes multiple field information, the above-mentioned row storage sub-table and the above-mentioned column storage sub-table inherit the multiple field information of the above-mentioned data parent table, and the above-mentioned row storage secondary sub-table and the above-mentioned column storage secondary sub-table inherit the multiple field information of the above-mentioned row storage sub-table.

[0070] Specifically, the above-mentioned data parent table also includes multiple field information, and the above-mentioned multiple field information can be understood as the attribute information of multiple attributes in the above-mentioned data parent table. For example, the above-mentioned field information can be age, gender, date of birth, etc., but is not limited to these exemplified field information. The row storage sub-table and the column storage sub-table inherit the multiple field information of the above-mentioned data parent table, and the row storage secondary sub-table and the column storage secondary sub-table inherit the multiple field information of the above-mentioned row storage sub-table. This ensures that the storage sub-tables obtained by subsequent splitting have the same field information as the data parent table, which ensures that data can be inserted and queried more conveniently, and further ensures high efficiency in querying and inserting data.

[0071] In one embodiment of the present application, the device further includes a first control unit and a second control unit, wherein the first control unit is used to control the data parent table to receive target request information, wherein the target request information is request information for inserting, deleting, changing or querying data; and the second control unit is used to control the data parent table to send the target request information to the corresponding storage sub-table according to the index information, and control the data parent table to send the response information of the storage sub-table to the target request information to the application. In this embodiment, according to the index information, the data parent table is controlled to send the target request information to the corresponding storage sub-table, that is, the corresponding storage sub-table is scanned and searched according to the index information, without querying other storage sub-tables, thus ensuring that the efficiency of scanning and searching the storage sub-table is high, and then the data parent table is controlled to send the response information of the storage sub-table to the target request information to the application, further ensuring that the occupied I / O resources are less.

[0072] It should be noted that the above storage sub-tables may include row storage sub-tables, column storage sub-tables, row storage secondary sub-tables, column storage secondary sub-tables, row storage tertiary sub-tables, column storage tertiary sub-tables, and subsequently split storage sub-tables.

[0073] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the following will be described in conjunction with specific embodiments:

[0074] Example

[0075] like Figure 4 and Figure 5 As shown, create a data parent table T1, create a row storage sub-table T1_P1 based on the data parent table, and use a time period greater than the current time period for sharding; create a column storage sub-table T1_P2 based on the data parent table, and use a time period less than or equal to the current time period for sharding; set the splitting cycle of the row storage sub-table in the database. After the splitting cycle is reached, the row storage sub-table T1_P1 is automatically split into a row storage sub-table T1P1 and a column storage sub-table T1P3. Therefore, the data parent table T1 contains three sub-tables, namely, the row storage sub-table T1P1, the column storage sub-table T1P2, and the column storage sub-table T1P3.

[0076] The processing device of the above-mentioned data storage structure includes a processor and a memory. The above-mentioned creation unit, the first determination unit and the first splitting unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.

[0077] The processor includes a kernel, which calls the corresponding program unit from the memory. One or more kernels can be set, and the problem of poor performance of inserting and querying data in HTAP scenarios in the prior art can be solved by adjusting kernel parameters.

[0078] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0079] An embodiment of the present invention provides a storage medium on which a program is stored. When the program is executed by a processor, the processing method of the data storage structure is implemented.

[0080] An embodiment of the present invention provides a processor, and the processor is used to run a program, wherein the processing method of the data storage structure is executed when the program is running.

[0081] In a typical embodiment of the present application, a data storage system is also provided, which includes: one or more processors, a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of the above methods.

[0082] The above-mentioned data storage system can execute any one of the above-mentioned methods. In the above-mentioned methods, first, a data parent table is created, and the above-mentioned data parent table includes index information for characterizing the mapping relationship between the above-mentioned data parent table and the row storage sub-table and the column storage sub-table, wherein the above-mentioned row storage sub-table is used to store data in the first target time period, and the above-mentioned column storage sub-table is used to store data before the above-mentioned first target time period; secondly, according to the splitting cycle, it is determined whether to split the above-mentioned row storage sub-table; finally, in the case of determining to split the above-mentioned row storage sub-table, the above-mentioned row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the above-mentioned row storage secondary sub-table is used to store data in the second target time period, and the above-mentioned column storage secondary sub-table is used to store data in the above-mentioned first target time period, and the above-mentioned second target time period is a time period after the above-mentioned first target time period. In this solution, the row storage subtable is split into a row storage secondary subtable and a column storage secondary subtable. The row storage secondary subtable is used to store data of the second target time period, that is, the row storage secondary subtable is used to store data close to the current time, and the column storage secondary subtable is used to store data far from the current time. This ensures that the obtained row storage secondary subtable is used to store the latest data, and the added column storage secondary subtable will not affect the hardware resources of the system. In addition, this solution implements the existence of both a row storage structure and a column storage structure in one table. Compared with the prior art that uses a multi-table merge method to process query data in the HTAP scenario, this solution performs queries in one table without the need to merge multiple tables. This not only ensures that less hardware resources are occupied and less I / O resources are consumed, but also ensures that the query efficiency is high by querying based on index information. In addition, when inserting or updating data, this solution does not need to write a logical process through an application program, which ensures that the efficiency of inserting and updating data is high, and also avoids the problem of high difficulty and high coupling in writing logical processes, thereby solving the problem of poor performance of inserting and querying data in the HTAP scenario in the prior art.

[0083] An embodiment of the present invention provides a device, the device including a processor, a memory, and a program stored in the memory and executable on the processor, and when the processor executes the program, at least the following steps are implemented:

[0084] Step S101, creating a data parent table, the data parent table including index information, the index information being used to characterize a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, the row storage sub-table being used to store data in a first target time period, and the column storage sub-table being used to store data before the first target time period;

[0085] Step S102, determining whether to split the row storage sub-table according to the splitting cycle;

[0086] Step S103, when it is determined that the row storage sub-table is to be split, the row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table is used to store data of the second target time period, and the column storage secondary sub-table is used to store data of the first target time period, and the second target time period is a time period after the first target time period.

[0087] The devices in this article can be servers, PCs, PADs, mobile phones, etc.

[0088] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program for initializing at least the following method steps:

[0089] Step S101, creating a data parent table, the data parent table including index information, the index information being used to characterize a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, the row storage sub-table being used to store data in a first target time period, and the column storage sub-table being used to store data before the first target time period;

[0090] Step S102, determining whether to split the row storage sub-table according to the splitting cycle;

[0091] Step S103, when it is determined that the row storage sub-table is to be split, the row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table is used to store data of the second target time period, and the column storage secondary sub-table is used to store data of the first target time period, and the second target time period is a time period after the first target time period.

[0092] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the above-mentioned units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0094] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0095] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0096] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the above-mentioned methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0097] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0098] 1) In the processing method of the data storage structure of the present application, firstly, a data parent table is created, and the data parent table includes index information for characterizing the mapping relationship between the data parent table and the row storage sub-table and the column storage sub-table, wherein the row storage sub-table is used to store data of a first target time period, and the column storage sub-table is used to store data before the first target time period; secondly, according to the splitting cycle, it is determined whether to split the row storage sub-table; finally, in the case of determining to split the row storage sub-table, the row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table is used to store data of a second target time period, and the column storage secondary sub-table is used to store data of the first target time period, and the second target time period is a time period after the first target time period. In this solution, the row storage subtable is split into a row storage secondary subtable and a column storage secondary subtable. The row storage secondary subtable is used to store data of the second target time period, that is, the row storage secondary subtable is used to store data close to the current time, and the column storage secondary subtable is used to store data far from the current time. This ensures that the obtained row storage secondary subtable is used to store the latest data, and the added column storage secondary subtable will not affect the hardware resources of the system. In addition, this solution implements the existence of both a row storage structure and a column storage structure in one table. Compared with the prior art that uses a multi-table merge method to process query data in the HTAP scenario, this solution performs queries in one table without the need to merge multiple tables. This not only ensures that less hardware resources are occupied and less I / O resources are consumed, but also ensures that the query efficiency is high by querying based on index information. In addition, when inserting or updating data, this solution does not need to write a logical process through an application program, which ensures that the efficiency of inserting and updating data is high, and also avoids the problem of high difficulty and high coupling in writing logical processes, thereby solving the problem of poor performance of inserting and querying data in the HTAP scenario in the prior art.

[0099] 2) In the processing device of the data storage structure of the present application, a creation unit is used to create a data parent table, the data parent table includes index information, the index information is used to characterize the mapping relationship between the data parent table and the row storage sub-table and the column storage sub-table, the row storage sub-table is used to store data of the first target time period, and the column storage sub-table is used to store data before the first target time period; the first determination unit is used to determine whether to split the row storage sub-table according to the splitting cycle; the first splitting unit is used to split the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table when it is determined that the row storage sub-table is split, the row storage secondary sub-table is used to store data of the second target time period, and the column storage secondary sub-table is used to store data of the first target time period, and the second target time period is a time period after the first target time period. In this solution, the row storage subtable is split into a row storage secondary subtable and a column storage secondary subtable. The row storage secondary subtable is used to store data of the second target time period, that is, the row storage secondary subtable is used to store data close to the current time, and the column storage secondary subtable is used to store data far from the current time. This ensures that the obtained row storage secondary subtable is used to store the latest data, and the added column storage secondary subtable will not affect the hardware resources of the system. In addition, this solution implements the existence of both a row storage structure and a column storage structure in one table. Compared with the prior art that uses a multi-table merge method to process query data in the HTAP scenario, this solution performs queries in one table without the need to merge multiple tables. This not only ensures that less hardware resources are occupied and less I / O resources are consumed, but also ensures that the query efficiency is high by querying based on index information. In addition, when inserting or updating data, this solution does not need to write a logical process through an application program, which ensures that the efficiency of inserting and updating data is high, and also avoids the problem of high difficulty and high coupling in writing logical processes, thereby solving the problem of poor performance of inserting and querying data in the HTAP scenario in the prior art.

[0100] 3) The data storage system of the present application can execute any of the above-mentioned methods. In the above-mentioned methods, first, a data parent table is created, and the above-mentioned data parent table includes index information for characterizing the mapping relationship between the above-mentioned data parent table and the row storage sub-table and the column storage sub-table, wherein the above-mentioned row storage sub-table is used to store data in the first target time period, and the above-mentioned column storage sub-table is used to store data before the above-mentioned first target time period; secondly, according to the splitting cycle, it is determined whether to split the above-mentioned row storage sub-table; finally, in the case of determining to split the above-mentioned row storage sub-table, the above-mentioned row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the above-mentioned row storage secondary sub-table is used to store data in the second target time period, and the above-mentioned column storage secondary sub-table is used to store data in the first target time period, and the above-mentioned second target time period is a time period after the above-mentioned first target time period. In this solution, the row storage subtable is split into a row storage secondary subtable and a column storage secondary subtable. The row storage secondary subtable is used to store data of the second target time period, that is, the row storage secondary subtable is used to store data close to the current time, and the column storage secondary subtable is used to store data far from the current time. This ensures that the obtained row storage secondary subtable is used to store the latest data, and the added column storage secondary subtable will not affect the hardware resources of the system. In addition, this solution implements the existence of both a row storage structure and a column storage structure in one table. Compared with the prior art that uses a multi-table merge method to process query data in the HTAP scenario, this solution performs queries in one table without the need to merge multiple tables. This not only ensures that less hardware resources are occupied and less I / O resources are consumed, but also ensures that the query efficiency is high by querying based on index information. In addition, when inserting or updating data, this solution does not need to write a logical process through an application program, which ensures that the efficiency of inserting and updating data is high, and also avoids the problem of high difficulty and high coupling in writing logical processes, thereby solving the problem of poor performance of inserting and querying data in the HTAP scenario in the prior art.

[0101] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for processing a data storage structure, characterized in that: include: Create a data parent table, the data parent table including index information, the index information is used to represent the mapping relationship between the data parent table and the row storage sub-table and the column storage sub-table, the row storage sub-table is used to store data in the first target time period, and the column storage sub-table is used to store data before the first target time period; Determining whether to split the row storage sub-table according to the splitting cycle; In the case where it is determined to split the row storage sub-table, split the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table, the row storage secondary sub-table is used to store data in a second target time period, and the column storage secondary sub-table is used to store data in the first target time period, and the second target time period is a time period after the first target time period; The data parent table includes at least one row storage sub-table and one column storage sub-table; In the case of determining to split the row storage sub-table, after splitting the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table, the method further includes: Determining whether to split the row storage secondary sub-table according to the split cycle; In the case of determining to split the row storage secondary sub-table, the row storage secondary sub-table is split into a row storage tertiary sub-table and a column storage tertiary sub-table, the row storage tertiary sub-table is used to store data of a third target time period, the column storage tertiary sub-table is used to store data of the second target time period, and the third target time period is a time period after the second target time period.

2. The method according to claim 1, characterized in that The row storage secondary subtable is used to process data in an OLTP scenario.

3. The method according to claim 1, characterized in that The column storage secondary subtable is used to process data in an OLAP scenario.

4. The method according to any one of claims 1 to 3, characterized in that: The data parent table also includes multiple field information, the row storage sub-table and the column storage sub-table inherit the multiple field information of the data parent table, and the row storage secondary sub-table and the column storage secondary sub-table inherit the multiple field information of the row storage sub-table.

5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Controlling the data parent table to receive target request information, wherein the target request information is request information for inserting, deleting, modifying or querying data; According to the index information, the data parent table is controlled to send the target request information to the corresponding storage sub-table, and the data parent table is controlled to send the response information of the storage sub-table to the target request information to the application.

6. A data storage structure processing device, characterized in that: include: a creating unit, configured to create a data parent table, the data parent table including index information, the index information being used to characterize a mapping relationship between the data parent table and a row storage sub-table and a column storage sub-table, the row storage sub-table being used to store data in a first target time period, and the column storage sub-table being used to store data before the first target time period; a first determining unit, configured to determine whether to split the row storage sub-table according to a splitting cycle; a first splitting unit, configured to, when it is determined that the row storage sub-table is to be split, split the row storage sub-table into a row storage secondary sub-table and a column storage secondary sub-table, wherein the row storage secondary sub-table is used to store data in a second target time period, and the column storage secondary sub-table is used to store data in the first target time period, wherein the second target time period is a time period after the first target time period; The data parent table includes at least one row storage sub-table and one column storage sub-table; In the case of determining that the row storage sub-table is to be split, after the row storage sub-table is split into a row storage secondary sub-table and a column storage secondary sub-table, the device further includes a second determination unit and a second splitting unit, wherein the second determination unit is used to determine whether to split the row storage secondary sub-table according to the splitting cycle; the second splitting unit is used to split the row storage secondary sub-table into a row storage tertiary sub-table and a column storage tertiary sub-table in the case of determining that the row storage secondary sub-table is to be split, wherein the row storage tertiary sub-table is used to store data of a third target time period, and the column storage tertiary sub-table is used to store data of the second target time period, and the third target time period is a time period after the second target time period.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method of any one of claims 1 to 5.

8. A processor, characterized in that: The processor is used to run a program, wherein the program executes the method according to any one of claims 1 to 5 when running.

9. A data storage system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 5.

Citation Information

Patent Citations

  • Hybrid database table stored as both row and column store

    CN103177056A

  • Data processing method and server

    CN108073584A