Data storage method and device based on row compression table, equipment and medium
By compressing the stored data before data storage and storing it in a row compression table, the limitations of traditional row storage in terms of storage efficiency and query performance are solved, and storage space utilization and data query performance are improved.
Patent Information
- Application Number
- CN202510652494.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional row storage has limitations in storage efficiency and query performance, resulting in low storage space utilization, especially in business scenarios with large data volume and large users, such as financial technology and medical health.
By determining the compression algorithm corresponding to the data type of data to be stored before data storage, the data to be stored is compressed, and the compressed data with a smaller volume is obtained and stored in the row compression table.
It effectively reduces the storage resources occupied by data in row storage, improves storage space utilization, and improves data query performance.
Smart Images

Figure CN120179658A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of database technology, and in particular, to a data storage method, apparatus, device, and medium based on a row-compressed table. Background Art
[0002] The efficient utilization of storage space and the optimization of data access performance are two core challenges in modern database management systems. With the exponential growth of data volume, traditional row storage gradually shows limitations in terms of storage efficiency and query performance. Row storage is suitable for point queries and transaction processing, but has a low storage space utilization rate.
[0003] For example, in fintech and healthcare business scenarios, as the number of users increases and the business develops, more and more data needs to be stored. Each user corresponds to a row of data. Storing data through traditional row storage requires a very large amount of storage resources and has a low storage space utilization rate. Summary of the Invention
[0004] Embodiments of this application provide a data storage method, apparatus, device, and medium based on a row-compressed table, which are used to reduce the storage resources occupied by data in row storage and improve the storage space utilization rate.
[0005] In a first aspect, embodiments of this application provide a data storage method based on a row-compressed table, which includes: In response to a storage instruction to store data to be stored in a row-compressed table, determine the data type of the data to be stored according to a preset data analysis algorithm; Determine a target compression algorithm corresponding to the data type from a preset plurality of compression algorithms; Compress the data to be stored according to the target compression algorithm to obtain compressed data, and add the algorithm identifier of the target compression algorithm to the metadata corresponding to the compressed data; Determine the data size of the compressed data, and determine the target storage location of the compressed data in the row-compressed table according to the data size; Store the compressed data at the target storage location.
[0006] In a second aspect, embodiments of this application also provide a data storage apparatus based on a row-compressed table, which includes: A transceiver unit, configured to obtain a storage instruction to store data to be stored in a row-compressed table; A processing unit, configured to determine the data type of the data to be stored according to a preset data analysis algorithm in response to the storage instruction; determine a target compression algorithm corresponding to the data type from a preset plurality of compression algorithms; perform compression processing on the data to be stored according to the target compression algorithm to obtain compressed data, and add an algorithm identifier of the target compression algorithm to metadata corresponding to the compressed data; determine the data size of the compressed data, and determine a target storage location of the compressed data in the row compression table according to the data size; store the compressed data at the target storage location.
[0007] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented.
[0008] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. The storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the above method can be implemented.
[0009] An embodiment of the present application provides a data storage method, apparatus, device, and medium based on a row compression table. Among them, the method includes: in response to a storage instruction to store data to be stored in a row compression table, determining the data type of the data to be stored according to a preset data analysis algorithm; determining a target compression algorithm corresponding to the data type from a preset plurality of compression algorithms; performing compression processing on the data to be stored according to the target compression algorithm to obtain compressed data, and adding an algorithm identifier of the target compression algorithm to metadata corresponding to the compressed data; determining the data size of the compressed data, and determining a target storage location of the compressed data in the row compression table according to the data size; storing the compressed data at the target storage location. An embodiment of the present application constructs a row compression table. Before storing data in the row compression table, it will first determine a compression algorithm corresponding to the data type of the data to be stored, and perform compression processing on the data to be stored through the corresponding compression algorithm to obtain compressed data with a smaller volume. Finally, the compressed data is stored in the row compression table, thereby effectively reducing the storage resources occupied by the data in row storage and improving the storage space utilization rate. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0011] Figure 1 Schematic structural diagram of the data storage system provided by an embodiment of the present application; Figure 2 Schematic flowchart of the data storage method based on a row compression table provided by an embodiment of the present application; Figure 3 Another schematic flowchart of the data storage method based on a row compression table provided by an embodiment of the present application; Figure 4 Schematic block diagram of the data storage device based on a row compression table provided by an embodiment of the present application; Figure 5 Schematic block diagram of the computer device provided by an embodiment of the present application. Detailed implementation manners
[0012] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0013] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0014] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0015] It should be further understood that the term " / and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0016] Embodiments of the present application provide a data storage method, device, equipment, and medium based on a row compression table.
[0017] The execution subject of the data storage method based on the row compression table can be the data storage device based on the row compression table provided in the embodiments of the present application, or a computer device integrated with the data storage device based on the row compression table. Among them, the data storage device based on the row compression table can be implemented in a hardware or software manner, and the computer device can be an electronic device with row storage functions such as a terminal or a server.
[0018] Among them, the data storage method based on the row compression table provided in the present application can be applied to business scenarios such as fintech, healthcare, and e-commerce.
[0019] In some embodiments, as Figure 1 shown, a data storage system is deployed in the computer device provided in the present application. In this embodiment, the data storage method based on the row compression table can be specifically executed through the data storage system. The data storage system can be a data storage system in the field of healthcare or fintech. For example, the data storage system of a financial database in financial business includes a data analysis module, a compression method selection module, a query location module, a decompression module, an index management module, a dynamic expansion module, and a maintenance tool module, where: The data analysis module is used to analyze the input data and determine the data type of each row of data; The compression method selection module is used to determine the compression algorithm for each row pair according to user selection or according to the data type; The query location module is used to locate the row data related to the row compression table according to the query condition (query key); The decompression module is used to decompress the located row data, restore the original data, and return and display it; The index management module is responsible for the creation, maintenance, and optimization of indexes, and supports efficient data query; The dynamic expansion module: is used to dynamically expand the storage space according to the change of the data volume of the row compression table; The maintenance tool module: is used to provide data storage and query operations.
[0020] The following takes the data storage system as the execution subject to elaborate in detail on the data storage method based on the row compression table provided in the present application. Among them, Figure 1 the structure of the data storage system shown is only for illustrative purposes. The data storage system provided in the present application can also be other structures, and the present application does not limit the specific structure of the data storage system.
[0021] Figure 2 is the flowchart of the data storage method based on the row compression table provided in the embodiments of the present application. As Figure 2 shown, the method includes the following steps S110 - S150.
[0022] S110. In response to a storage instruction for storing data to be stored into a row-compressed table, determine the data type of the data to be stored according to a preset data analysis algorithm.
[0023] In this embodiment, when a user needs to store data, a storage instruction will be sent to the data storage system, so that the data storage system automatically stores the data to be stored indicated by the storage instruction into the row-compressed table. Wherein, the data to be stored in this embodiment is a row of data. The user can store multiple rows of data to be stored into the row-compressed table at the same time. At this time, the data storage system sequentially executes steps S110 - S150 for each row of data to be stored.
[0024] In some embodiments, the preset data analysis algorithm includes numerical type detection, text type detection, and text length detection. The obtained data types include large numerical type, small numerical type, large text type, and small text type. Among them, the large numerical type is to detect that the data to be stored is numerical data and the data length is greater than or equal to the preset length. The small numerical type is to detect that the data to be stored is numerical data and the data length is less than the preset length. The large text type is to detect that the data to be stored is text data and the data length is greater than or equal to the preset length. The small text type is to detect that the data to be stored is text data and the data length is less than the preset length. Among them, in practical applications, other types of data analysis algorithms can also be set according to needs. This embodiment does not limit the specific type of the data analysis algorithm.
[0025] For example, in a financial business scenario, currently, it is necessary to store the personal business file of a customer A, which includes information such as customer ID, customer name, customer gender, customer phone number, customer address, and types of financial products purchased. The user inputs the personal business file through the data storage interface provided by the data storage system and triggers a storage operation. At this time, the data storage system automatically determines the data type corresponding to the personal business file through a preset data analysis algorithm. For example, it is analyzed that the personal business file is of the small text type.
[0026] Another example is that in the scenario of storing medical records in the medical and health business, after obtaining the medical records of a customer, the data storage method based on the row-compressed table provided by this application can be used for data storage of medical records.
[0027] S120. Determine a target compression algorithm corresponding to the data type from a preset multiple compression algorithms.
[0028] In the data storage system provided in this embodiment, multiple different types of compression algorithms are preset, and a corresponding relationship between the data type and the compression algorithm is preset. After obtaining the data type of the data pair to be stored, the target compression algorithm will be determined from multiple different types of compression algorithms based on the corresponding relationship between the data type and the compression algorithm and the data type of the data pair to be stored.
[0029] For example, for large numerical types and large text types, the corresponding compression algorithm is set to the zstd compression algorithm to provide a high compression ratio. For small numerical types and small text types, the corresponding compression algorithm is the lz4 compression algorithm to improve the compression speed.
[0030] For example, in the financial business scenario, currently it is necessary to store the personal business file A of customer A and the personal business file B of customer B. By analysis, it is determined that the data type corresponding to the personal business file A is a small text type, and the data type corresponding to the personal business file B is a large text type. At this time, according to the corresponding relationship, the compression algorithm corresponding to the personal business file A is determined to be the lz4 compression algorithm, and the compression algorithm corresponding to the personal business file B is determined to be the zstd compression algorithm.
[0031] In addition, in some embodiments, the present application also supports manual selection of the compression algorithm, that is, after the user views the data type of the data to be stored, the user independently selects the currently required compression algorithm from multiple compression algorithms as the target compression algorithm.
[0032] For example, in the financial business scenario, currently it is necessary to quickly store the personal business file B of customer B. Although the data type of the personal business file B is identified as a large text type at this time, the user needs to store this data as soon as possible, so the user manually selects the lz4 compression algorithm with a faster storage speed.
[0033] S130. Compress the data to be stored according to the target compression algorithm to obtain compressed data, and add the algorithm identifier of the target compression algorithm to the metadata corresponding to the compressed data.
[0034] In this embodiment, after determining the target compression algorithm, the data to be stored is compressed to obtain compressed data. Since different types of data use different compression algorithms, in this embodiment, the algorithm identifier of the target compression algorithm is also added to the metadata corresponding to the compressed data to facilitate subsequent decompression operations.
[0035] Among them, the metadata corresponding to the compressed data includes not only the algorithm identifier, but also the row ID, the number of columns, the original data length, the compressed length, and the number of blocks, etc. The compressed length is used for the verification of the compressed data, the original data length is used for the verification of the subsequent decompressed data, and the number of blocks is used for the data block splicing verification during subsequent decompression.
[0036] S140. Determine the data size of the compressed data, and determine the target storage location of the compressed data in the row compression table according to the data size.
[0037] In this embodiment, after determining the data size of the compressed data, a storage space of a corresponding size can be allocated for the compressed data, and the target storage location of the compressed data in the row compression table is determined.
[0038] S150. Store the compressed data at the target storage location.
[0039] In this embodiment, after compressing the data and determining the target storage location corresponding to the compressed data, the compressed data is stored at the position corresponding to the target storage location in the row compression table.
[0040] In some embodiments, to improve the retrieval speed, this embodiment also constructs a table index for the row compression table. At this time, before performing the compression process on the data to be stored according to the target compression algorithm to obtain the compressed data, the method further includes: extracting a target index key from the data to be stored according to a preset first index key extraction rule; In this embodiment, before compressing, a target index key is first extracted from the data to be stored based on the first index key extraction rule. The index construction of the current row does not require decompressing the corresponding data, which can improve the index construction efficiency. Among them, the first index key extraction rule can be to use the information of a specified column as the target index key. For example, the personal business file includes customer ID, customer name, customer gender, customer phone, customer address, and the type of financial product purchased. The information corresponding to the customer ID is used as the target index key. For example, if the customer ID is 00021, the target index key is 00021.
[0041] It should be noted that this application involves data related to personal information (such as customer ID). When applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0042] At this time, after storing the compressed data at the target storage location, the method further includes: in response to an index update instruction of the row compression table, updating the table index corresponding to the row compression table according to the target index key and the target storage location.
[0043] In this embodiment, the index update instruction is automatically triggered by the system or manually triggered after data storage. Specifically, the index type of the table index is a b-tree index, a hash index, a gist index, or a brin index. The row compression table in this embodiment is constructed based on a target database, and the target database is a PG database (PostgreSQL).
[0044] In some embodiments, in response to an index update instruction for the row compression table, updating the table index corresponding to the row compression table according to the target index key and the target storage location includes: In response to the index update instruction for the row compression table, determine whether there is a table index corresponding to the row compression table; if there is the table index, update the table index according to the target index key and the target storage location.
[0045] In this embodiment, it is possible that there is no index table built for the row compression table before. At this time, when the user wants to build an index for the row compression table, an index update instruction of the system can be triggered. After obtaining the index update instruction for the row compression table, it will be checked whether there is a table index corresponding to the row compression table. If there is, the table index will be updated.
[0046] If there is no such table index, scan each tuple in the row compression table, perform a predicate evaluation on each tuple, and determine the index key type according to the predicate evaluation result; extract the index key corresponding to each tuple according to the index key type; construct the table index according to the index key corresponding to each tuple and the storage location of each tuple in the row compression table.
[0047] Among them, to ensure the legality of index construction and environment preparation, before scanning, first determine whether each tuple supports partial range scanning (that is, check whether there is a predicate that meets the preset search conditions, etc.), and report an error for the rows that do not support partial range scanning. And perform execution state and expression context preparation: create an execution state (EState) and an expression context (ExprContext) for processing index expressions and predicates.
[0048] The above tuples are decompressed compressed data. By decompressing the compressed data and scanning each tuple in the row compression table, the predicates of each tuple (such as the attributes of each column) can be obtained, and then the predicates of each tuple are evaluated according to the predicates of each tuple to determine the predicate suitable as the index key. For example, the column type with the most complete data volume in the corresponding column can be used as the index key type, such as using the customer ID as the index key type.
[0049] Further, in some embodiments, scanning each tuple in the row compression table includes: obtaining the current operation state corresponding to the row compression table; if the operation state is a serial operation state, obtaining a private snapshot corresponding to the row compression table, and scanning each tuple in the row compression table based on the private snapshot, where the private snapshot is a snapshot obtained when the index update instruction is triggered; if the operation state is a parallel operation state, obtaining a shared snapshot corresponding to the row compression table, and scanning each tuple in the row compression table based on the shared snapshot, where the shared snapshot is a snapshot obtained when the row compression table is scanned.
[0050] Specifically, in this embodiment, the row compression table is scanned through a snapshot, which can ensure data consistency.
[0051] In some embodiments, referring to Figure 3 , after updating the table index corresponding to the row compression table according to the target index key and the target storage location in response to the index update instruction of the row compression table, the method further includes steps S160 - S1100: S160. Obtain a data query instruction for the row compression table, where the data query instruction carries a query key; S170. Determine whether the query key type of the query key is the same as the index key type of the table index; S180. If the query key type is the same as the index key type, determine target compressed data in the row compression table according to the query key and the table index; S190. Decompress the target compressed data according to the algorithm identifier in the metadata of the target compressed data to obtain target query data; S1100. If the query key type is different from the index key type, perform a forward scan on the row compression table according to the query key to determine the target query data from the row compression table.
[0052] Specifically, if the index key type of the table index is the customer ID, and the query key type corresponding to the query key is also the customer ID, it indicates that the query key type is the same as the index key type; otherwise, the query key type (for example, the query key type is the customer phone) is different from the index key type.
[0053] Among them, determining the target compressed data in the row compressed table according to the query key and the table index includes: determining the corresponding storage location from the table index according to the query key, and then determining the compressed data at the corresponding storage location in the row compressed table as the target compressed data. Scanning the row compressed table forward according to the query key to determine the target query data from the row compressed table includes: decompressing the compressed data of the row compressed table sequentially according to the query key until the row data corresponding to the query key is found from the decompressed data as the target query data.
[0054] Further, the target compressed data includes a plurality of target data blocks; decompressing the target compressed data according to the algorithm identifier in the metadata of the target compressed data to obtain the target query data includes: splicing the plurality of target data blocks to obtain a target data block group; decompressing the target data block group according to the algorithm identifier in the metadata of the target compressed data to obtain the target query data.
[0055] This embodiment can splice multiple data blocks and then decompress them, thereby improving the data decompression speed and the data query speed.
[0056] In some embodiments, before compressing the data to be stored according to the target compression algorithm to obtain the compressed data, the method further includes: extracting a plurality of target index keys from the data to be stored according to a preset second index key extraction rule; after compressing the data to be stored according to the target compression algorithm to obtain the compressed data, the method further includes: constructing a Bloom filter according to the plurality of target index keys and adding the Bloom filter to the metadata corresponding to the compressed data; after storing the compressed data in the target storage location, the method further includes: obtaining a data query instruction for the row compressed table, where the data query instruction carries a query key; determining the Bloom filter in which the query key exists among the plurality of Bloom filters of the row compressed table as the target Bloom filter; determining the compressed data corresponding to the target Bloom filter as the target compressed data; decompressing the target compressed data according to the algorithm identifier in the metadata of the target compressed data to obtain the target query data.
[0057] This embodiment can quickly filter the compressed data through the Bloom filter, which can improve the data query speed. In addition, the Bloom filter has the characteristics of small volume and small occupied storage space, and multiple index keys can be stored in the Bloom filter, enabling users to quickly retrieve data using various types of query keys.
[0058] In some embodiments, after constructing the table index according to the index keys respectively corresponding to each tuple and the storage positions of each tuple in the row compression table, the method further includes: Obtain a data query instruction for the row compression table, where the data query instruction carries a query key; perform a forward scan on the row compression table according to the query key to determine target query data from the row compression table.
[0059] In this embodiment, query conditions such as indexes may not be set. After obtaining the query key, directly perform a forward scan on the row compression table, decompress the data in the row compression table in sequence until the data corresponding to the query key is scanned as the target query data.
[0060] For example, in a financial business scenario, a user inputs "Customer ID: 00021" into the data storage system. At this time, the system automatically scans from the row compression database to obtain the target query data corresponding to "Customer ID: 00021".
[0061] Furthermore, this embodiment also supports dynamically expanding the storage space of the row compression table. Specifically, obtain the total number of bytes currently occupied by the row compression table at a preset time interval (such as 10 minutes), and determine the remaining space according to the occupied total number of bytes and the total space size. If the remaining space is less than the preset space threshold, automatically apply for new space to achieve dynamic expansion of the storage space.
[0062] In summary, in this embodiment, in response to a storage instruction to store data to be stored into the row compression table, determine the data type of the data to be stored according to a preset data analysis algorithm; determine a target compression algorithm corresponding to the data type from a preset plurality of compression algorithms; perform compression processing on the data to be stored according to the target compression algorithm to obtain compressed data, and add the algorithm identifier of the target compression algorithm to the metadata corresponding to the compressed data; determine the data size of the compressed data, and determine the target storage position of the compressed data in the row compression table according to the data size; store the compressed data to the target storage position. The embodiment of the present application constructs a row compression table. Before storing data into the row compression table, it will first determine the compression algorithm corresponding to the data type of the data to be stored, and perform compression processing on the data to be stored through the corresponding compression algorithm to obtain compressed data with a smaller volume. Finally, store the compressed data into the row compression table, thereby effectively reducing the storage resources occupied by data in row storage and improving the storage space utilization rate.
[0063] Figure 4 It is a schematic block diagram of a data storage device 400 based on a row compression table provided by an embodiment of the present application. As Figure 4As shown in the figure, corresponding to the above data storage method based on the row compression table, the present application also provides a data storage device 400 based on the row compression table. The data storage device 400 based on the row compression table includes a unit for executing the above data storage method based on the row compression table, and the data storage device 400 based on the row compression table can be configured in a terminal or a server. Specifically, please refer to Figure 4 , the data storage device 400 based on the row compression table includes a transceiver unit 401 and a processing unit 402, where: The transceiver unit 401 is used to obtain a storage instruction for storing the data to be stored into the row compression table; The processing unit 402 is used to, in response to the storage instruction, determine the data type of the data to be stored according to a preset data analysis algorithm; determine a target compression algorithm corresponding to the data type from a preset plurality of compression algorithms; perform compression processing on the data to be stored according to the target compression algorithm to obtain compressed data, and add an algorithm identifier of the target compression algorithm to metadata corresponding to the compressed data; determine the data size of the compressed data, and determine a target storage location of the compressed data in the row compression table according to the data size; store the compressed data at the target storage location.
[0064] In some embodiments, before the processing unit 402 executes the step of performing compression processing on the data to be stored according to the target compression algorithm to obtain compressed data, it is further used to: Extract a target index key from the data to be stored according to a preset first index key extraction rule; After processing the step of storing the compressed data at the target storage location, it is further used to: In response to an index update instruction of the row compression table, update the table index corresponding to the row compression table according to the target index key and the target storage location.
[0065] In some embodiments, when the processing unit 402 executes the step of updating the table index corresponding to the row compression table according to the target index key and the target storage location in response to the index update instruction of the row compression table, it specifically uses: In response to the index update instruction of the row compression table, determine whether there is a table index corresponding to the row compression table; if there is the table index, update the table index according to the target index key and the target storage location; At this time, after the processing unit 402 executes the step of determining whether there is a table index corresponding to the row compression table, it is further used to: If the table index does not exist, scan each tuple in the row-compressed table, perform predicate evaluation on each tuple, determine the index key type according to the predicate evaluation result; extract the index key corresponding to each tuple according to the index key type; construct the table index according to the index key corresponding to each tuple and the storage location of each tuple in the row-compressed table.
[0066] In some embodiments, when the processing unit 402 executes the step of scanning each tuple in the row-compressed table, it is specifically configured to: Obtain the operation state currently corresponding to the row-compressed table; if the operation state is a serial operation state, obtain the private snapshot corresponding to the row-compressed table, and scan each tuple in the row-compressed table based on the private snapshot, where the private snapshot is the snapshot obtained when the index update instruction is triggered; if the operation state is a parallel operation state, obtain the shared snapshot corresponding to the row-compressed table, and scan each tuple in the row-compressed table based on the shared snapshot, where the shared snapshot is the snapshot obtained when the row-compressed table is scanned.
[0067] In some embodiments, after the processing unit 402 executes the step of updating the table index corresponding to the row-compressed table according to the target index key and the target storage location in response to the index update instruction of the row-compressed table, it is further configured to: Obtain a data query instruction for the row-compressed table through the transceiver unit 401, where the data query instruction carries a query key; determine whether the query key type of the query key is the same as the index key type of the table index; if the query key type is the same as the index key type, determine the target compressed data in the row-compressed table according to the query key and the table index; perform decompression processing on the target compressed data according to the algorithm identifier in the metadata of the target compressed data to obtain the target query data; if the query key type is different from the index key type, perform a forward scan on the row-compressed table according to the query key to determine the target query data from the row-compressed table.
[0068] In some embodiments, the target compressed data includes a plurality of target data blocks; when the processing unit 402 executes the step of performing decompression processing on the target compressed data according to the algorithm identifier in the metadata of the target compressed data to obtain the target query data, it is specifically configured to: Perform splicing processing on the plurality of target data blocks to obtain a target data block group; perform decompression processing on the target data block group according to the algorithm identifier in the metadata of the target compressed data to obtain the target query data.
[0069] In some embodiments, before the processing unit 402 executes the step of compressing the data to be stored according to the target compression algorithm to obtain compressed data, it is further configured to: Extract a plurality of target index keys from the data to be stored according to a preset second index key extraction rule; After executing the step of compressing the data to be stored according to the target compression algorithm to obtain compressed data, it is further configured to: Construct a Bloom filter according to the plurality of target index keys, and add the Bloom filter to the metadata corresponding to the compressed data; After executing the step of storing the compressed data in the target storage location, it is further configured to: Obtain a data query instruction for the row compression table through the transceiver unit 401, where the data query instruction carries a query key; determine the Bloom filter in which the query key exists among the plurality of Bloom filters of the row compression table as the target Bloom filter; determine the compressed data corresponding to the target Bloom filter as the target compressed data; and decompress the target compressed data according to the algorithm identifier in the metadata of the target compressed data to obtain target query data.
[0070] In summary, in this embodiment, through the data storage device 400 based on the row compression table, before storing the data in the row compression table, the compression algorithm corresponding to the data type of the data to be stored will be determined first, and the data to be stored will be compressed through the corresponding compression algorithm to obtain compressed data with a smaller volume. Finally, the compressed data will be stored in the row compression table, thereby effectively reducing the storage resources occupied by the data in the row storage and improving the storage space utilization rate.
[0071] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above data storage device based on the row compression table and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and conciseness of description, they will not be elaborated here.
[0072] The above data storage device based on the row compression table can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 5 shown.
[0073] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a terminal or a server.
[0074] Refer to Figure 5, the computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.
[0075] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which when executed, can cause the processor 502 to execute a data storage method based on a row compression table.
[0076] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0077] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, it can cause the processor 502 to execute a data storage method based on a row compression table.
[0078] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0079] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps: In response to a storage instruction to store the data to be stored into the row compression table, determine the data type of the data to be stored according to a preset data analysis algorithm; Determine the target compression algorithm corresponding to the data type from a preset plurality of compression algorithms; Compress the data to be stored according to the target compression algorithm to obtain compressed data, and add the algorithm identifier of the target compression algorithm to the metadata corresponding to the compressed data; Determine the data size of the compressed data, and determine the target storage location of the compressed data in the row compression table according to the data size; Store the compressed data at the target storage location.
[0080] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0081] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0082] Therefore, the present application also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps: In response to a storage instruction for storing data to be stored into a row compression table, determine the data type of the data to be stored according to a preset data analysis algorithm; Determine a target compression algorithm corresponding to the data type from a preset plurality of compression algorithms; Perform compression processing on the data to be stored according to the target compression algorithm to obtain compressed data, and add an algorithm identifier of the target compression algorithm to the metadata corresponding to the compressed data; Determine the data size of the compressed data, and determine a target storage location of the compressed data in the row compression table according to the data size; Store the compressed data at the target storage location.
[0083] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, an optical disk, or other computer-readable storage media that can store program codes.
[0084] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0085] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0086] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of this application can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0087] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application.
[0088] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A data storage method based on a row compression table, characterized in that: include: In response to a storage instruction to store the data to be stored in the row compression table, determining the data type of the data to be stored according to a preset data analysis algorithm; Determining a target compression algorithm corresponding to the data type from a plurality of preset compression algorithms; Compressing the data to be stored according to the target compression algorithm to obtain compressed data, and adding the algorithm identifier of the target compression algorithm to the metadata corresponding to the compressed data; Determining a data size of the compressed data, and determining a target storage location of the compressed data in the row compression table according to the data size; The compressed data is stored in the target storage location.
2. The method according to claim 1, characterized in that Before compressing the data to be stored according to the target compression algorithm to obtain compressed data, the method further includes: Extracting a target index key from the data to be stored according to a preset first index key extraction rule; After storing the compressed data in the target storage location, the method further includes: In response to the index update instruction of the row compression table, a table index corresponding to the row compression table is updated according to the target index key and the target storage location.
3. The method according to claim 2, characterized in that The step of responding to the index update instruction of the row compression table and updating the table index corresponding to the row compression table according to the target index key and the target storage location includes: In response to an index update instruction of the row compression table, determining whether there is a table index corresponding to the row compression table; If the table index exists, updating the table index according to the target index key and the target storage location; The method further comprises: If the table index does not exist, scanning each tuple in the row compression table, performing predicate evaluation on each tuple, and determining the index key type according to the predicate evaluation result; Extract the index keys corresponding to each tuple according to the index key type; The table index is constructed according to the index keys corresponding to the tuples and the storage positions of the tuples in the row compression table.
4. The method according to claim 3, characterized in that The scanning of each tuple in the row compression table includes: Obtaining the current operation status corresponding to the row compression table; If the operation state is a serial operation state, a private snapshot corresponding to the row compression table is obtained, and each tuple in the row compression table is scanned based on the private snapshot, where the private snapshot is a snapshot obtained when the index update instruction is triggered; If the operation state is a parallel operation state, a shared snapshot corresponding to the row compression table is obtained, and each tuple in the row compression table is scanned based on the shared snapshot, where the shared snapshot is a snapshot obtained when the scan of the row compression table is triggered.
5. The method according to claim 2, characterized in that: After responding to the index update instruction of the row compression table and updating the table index corresponding to the row compression table according to the target index key and the target storage location, the method further includes: Obtaining a data query instruction for the row compression table, wherein the data query instruction carries a query key; Determining whether the query key type of the query key is the same as the index key type of the table index; If the query key type is the same as the index key type, determining target compressed data in the row compression table according to the query key and the table index; Decompressing the target compressed data according to the algorithm identifier in the metadata of the target compressed data to obtain target query data; If the query key type is different from the index key type, the row compression table is forward scanned according to the query key to determine the target query data from the row compression table.
6. The method according to claim 2, characterized in that The target compressed data includes a plurality of target data blocks; the target compressed data is decompressed according to the algorithm identifier in the metadata of the target compressed data to obtain the target query data, including: Splicing a plurality of the target data blocks to obtain a target data block group; The target data block group is decompressed according to the algorithm identifier in the metadata of the target compressed data to obtain target query data.
7. The method according to claim 1, characterized in that Before compressing the data to be stored according to the target compression algorithm to obtain compressed data, the method further includes: Extracting a plurality of target index keys from the data to be stored according to a preset second index key extraction rule; After compressing the data to be stored according to the target compression algorithm to obtain compressed data, the method further includes: Constructing a Bloom filter according to the plurality of target index keys, and adding the Bloom filter to metadata corresponding to the compressed data; After storing the compressed data in the target storage location, the method further includes: Obtaining a data query instruction for the row compression table, wherein the data query instruction carries a query key; Determine a Bloom filter storing the query key among the multiple Bloom filters of the row compression table as a target Bloom filter; Determining the compressed data corresponding to the target Bloom filter as target compressed data; The target compressed data is decompressed according to the algorithm identifier in the metadata of the target compressed data to obtain target query data.
8. A data storage device based on a row compression table, characterized in that: include: A transceiver unit, used for obtaining a storage instruction for storing the data to be stored in the row compression table; A processing unit, configured to determine the data type of the data to be stored according to a preset data analysis algorithm in response to the storage instruction; and determine a target compression algorithm corresponding to the data type from a plurality of preset compression algorithms; Compressing the data to be stored according to the target compression algorithm to obtain compressed data, and adding the algorithm identifier of the target compression algorithm to the metadata corresponding to the compressed data; Determine the data size of the compressed data, and determine the target storage location of the compressed data in the row compression table according to the data size; and store the compressed data in the target storage location.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the data storage method based on the row compression table according to any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the data storage method based on a row compression table according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data Compression For Reducing Storage Requirements In A Database System
CN102804168A
Distributed type in-memory database indexing method oriented to structural data
CN105095520A
Database management method, device and electronic equipment
CN105117489A
Data processing method and apparatus, storage medium, and electronic apparatus
CN109241102A
Method and system for recovering MSSQL database and storage medium
CN113282592A