Data intelligent compression method and system for embedded database

Through intelligent data compression methods, adapting to different environments of embedded databases, the problem of slow compression speed is solved, efficient data compression and decompression is achieved, and resource utilization and network transmission speed are improved.

CN116841973BActive Publication Date: 2025-06-03HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310830705.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-07
Publication Date
2025-06-03
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

Embedded databases cannot be adapted to different environments, resulting in slow compression speed.

Method used

Using an intelligent data compression method for embedded databases, we use CPU, memory and hard disk usage to determine whether there is an embedded database connection, use transactions to statistical tables, use K-Means algorithm to classify data, evaluate data collection characteristics, select appropriate compression algorithms, and optimize compression algorithms through Q learning.

Benefits of technology

It improves the compression speed of embedded databases in different environments, realizes intelligent judgment of compression timing, saves storage resource space, improves resource utilization, and accelerates network transmission speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116841973B_ABST
    Figure CN116841973B_ABST
Patent Text Reader

Abstract

An intelligent data compression method and system for embedded databases, which relate to the field of data compression technology. Aiming at the problem in the prior art that embedded databases cannot adapt to different environments, resulting in slow compression speed, this application can classify and identify different scenarios and system conditions, and then automatically select the required compression algorithm and automatically decompress when needed. It can adapt to a variety of usage environments, improve the compression speed of embedded databases in different environments, and can intelligently judge the compression timing, such as the device environments of Internet of Things devices, mobile phones, personal computers, etc., and make adaptive adjustments according to the resource quantities in these environments to achieve intelligent compression and decompression of embedded databases, ultimately achieving the purpose of saving storage resource space, improving resource utilization rate, and accelerating network transmission speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data compression, and specifically to an intelligent data compression method and system for an embedded database. Background Art

[0002] Data compression technology is one of the current research focuses of databases. The significance of researching data compression algorithms lies in: firstly, the compressed data can improve the utilization rate of storage space, thereby reducing the hardware cost of storing data; secondly, when transmitting the same amount of data over a network, the number of bytes transmitted can be reduced in the compressed state, which can greatly improve the throughput rate of data transmission and, in effect, increase the network transmission bandwidth. These two aspects can save a large amount of hardware and software costs for enterprises or individuals. Due to the urgent need of enterprises to reduce hardware costs, the current exploration of data compression technology mainly focuses on data compression in server-side databases. For example, the gzip algorithm improved based on the LZ4 algorithm and the ZSTD algorithm is widely used. These general and relatively cumbersome compression algorithms are called heavy algorithms. On the other hand, the relatively light algorithms, including the delta-delta algorithm, RLE (run-length encoding), and other compression algorithms specifically for bit operations, are also often used. In the academic field, the recent hot topic is to integrate neural network or reinforcement learning algorithms into compression algorithms to make them more intelligent. For example, compression algorithms improved using VAE, lossy compression algorithms improved using GAN, etc., have all achieved good results.

[0003] For existing algorithms, whether they are heavy algorithms or light algorithms, whether in the industrial field or the academic field, the database backgrounds they mainly study, classified by application mode, focus on distributed databases; from the perspective of storage mode, they mainly focus on databases with row storage and column storage; from the different types of stored data, they mainly focus on relational databases and time-series databases. All of the above application scenarios have one thing in common, that is, these algorithms are often deployed on servers to provide solutions specifically for data compression in databases on servers. The characteristic of databases on servers is that a single database runs on a single device, and the resources of the device are often concentrated on the database for its use. Therefore, the algorithms proposed so far are often implemented on the premise of sufficient resources (especially computing resources), a reliable network and device environment, and less consideration of device factors such as power consumption. However, for embedded databases, their operating environments are often relatively harsh. They may not only be deployed in server computer rooms but also in environments with large interference and instability, such as factories and farmlands. In such environments, the network resources and computing resources available to embedded databases are both reduced, and their power consumption issues often need to be considered to prevent power loss or cost increase due to excessive power consumption. The complex environment and structure of the deployment of embedded databases determine that they require a specially customized compression algorithm.

[0004] On the other hand, embedded databases are widely used as databases that have been downloaded 50 billion times so far, including Internet of Things devices (such as temperature sensors, monitoring devices, etc.), mobile phones (a large number of embedded databases are used in both Android and iOS), personal computers (including both programs and system services), etc. In this case, it is almost impossible for a DBA to deeply optimize the database performance by tracking the code in real time. After one deployment, the embedded database can only mechanically keep running. This results in the inability of the embedded database to adapt to different environments, leading to slow compression speed. Summary of the Invention

[0005] The object of the present invention is to propose a data intelligent compression method and system for embedded databases in view of the problem that in the prior art, embedded databases cannot adapt to different environments, resulting in slow compression speed.

[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows:

[0007] A data intelligent compression method for embedded databases includes the following steps:

[0008] Step 1: Detect the current usage of CPU, memory, and hard disk. If the usage rate of any one of the CPU and memory exceeds 70%, no compression is performed. If the usage rates of both the CPU and memory do not exceed 70%, compression is performed, and step 2 is executed;

[0009] Step 2: Determine whether there is an embedded database connection currently. If there is no embedded database connection currently, connect to the embedded database. If there is an embedded database connection currently, send information to each embedded database to inquire about the current transaction status and restrictions of the embedded database. For an embedded database that is in the middle of a transaction or has data modification restrictions, no compression is performed. Otherwise, compression is performed, and step 3 is executed;

[0010] Step 3: For the embedded database to be compressed, first count the most recent usage transactions of each table in the embedded database. If there is data in the table that has not been accessed within more than 100 data-related transactions, the table is marked as an infrequently used table. Then, use the K-Means algorithm to classify the data in each infrequently used table by primary key or timestamp, and divide each table into different data sets;

[0011] Step 4: Perform feature evaluation at the data level on each data set obtained in step 3, and select the data sets that need to be compressed according to the evaluation results;

[0012] The specific content of step 4 is as follows:

[0013] Set the read operation transaction metric value of the previous embedded database to increase by 1 for every 10 unread transactions, set the write operation transaction metric value of the previous embedded database to increase by 1 for every 10 unwritten transactions, record the metric value of the regular structure in the data itself as 1, and the metric value of the irregular structure as 0. Set the weight of the read operation to 1, the weight of the write operation to 2, and the structure weight of the data itself to 1. Multiply all metric values by their weights to obtain the total weight value of a data set. Sort the total weight values of all data sets from high to low, and select the data set with the highest total weight value for data compression;

[0014] Step Five: Use Q-learning as a way to select a compression algorithm. Take the current CPU, memory, and hard disk usage as the state input of Q-learning, and the selected compression algorithm as the output. Calculate the reward using the throughput improvement of the compressed embedded database and the reduction of occupied resources. Obtain the compression algorithm through each iteration;

[0015] Step Six: Compress the data set obtained in Step Four using the compression algorithm selected in Step Five.

[0016] Further, the specific steps of Step One are as follows:

[0017] Create a daemon process, and use the daemon process to detect the current CPU, memory, and hard disk usage. If the usage rate of any one of the CPU and memory exceeds 70%, do not perform compression, and the daemon process enters the sleep state, waiting to be woken up again for judgment. If the usage rates of both the CPU and memory do not exceed 70%, perform compression and execute Step Two.

[0018] Further, the specific steps of creating the daemon process are as follows:

[0019] Use a python script to create a daemon process that wakes up every 5 minutes. This daemon process will detect the current CPU, memory, and hard disk usage every 5 minutes.

[0020] Further, the specific steps of connecting to the embedded database are as follows:

[0021] Send a connection request to each embedded database through the port number. After the embedded database agrees to the connection, connect to the database and ensure the connection of the embedded database through the heartbeat mechanism during subsequent operation.

[0022] Further, the maximum number of data rows in the set in Step Three is 1000 rows.

[0023] Further, the specific way of using Q-learning as a way to select a compression algorithm in Step Five is as follows:

[0024] When the CPU or memory usage rate reaches more than 50% and less than 70%, a lightweight compression algorithm is selected;

[0025] When both the CPU and memory usage rates are less than 50% and the hard disk usage rate is less than 50%, a lightweight compression algorithm is selected. If the hard disk usage rate is not less than 50%, a heavyweight compression algorithm is selected;

[0026] When the hard disk usage rate reaches more than 95%, at this time, the CPU and memory usage rates will be ignored, and the heavyweight algorithm will be directly used to quickly compress the data in the hard disk.

[0027] Furthermore, the lightweight compression algorithm includes RLE or delta-of-delta, and the heavyweight compression algorithm is the Gzip algorithm.

[0028] Furthermore, the specific steps of compression in step six are as follows:

[0029] When compressing, first extract the data set to be compressed from the corresponding table in the embedded database, store the extracted data set in a separate file, use the compression algorithm selected in step five to compress this file, obtain the compressed file, and finally store the compressed file in an additional folder, and all compressed data files are saved in this folder.

[0030] Furthermore, the method further includes step seven, and the specific steps of step seven are as follows:

[0031] Record all the compressed data files saved in the folder, and use a B+ tree to establish an index for the data files. At the same time, monitor the transactions of the embedded database in real time. Once there is a transaction involving the compressed data file, find the compressed data file through the index, decompress the data file, and then merge the decompressed data back into the table of the embedded database, and finally update the read or write transaction records of the corresponding data.

[0032] An intelligent data compression system for an embedded database, the system includes a system detection module, a connection judgment module, a data set classification module, a data set evaluation module, and a Q learning module;

[0033] The system detection module is used to detect the current CPU, memory, and hard disk usage conditions. If the usage rate of any one of the CPU and memory exceeds 70%, no compression is performed. If the usage rates of both the CPU and memory do not exceed 70%, compression is performed;

[0034] The connection judgment module is used to judge whether there is an embedded database connection currently. If there is no embedded database connection currently, it connects to the embedded database. If there is an embedded database connection currently, it sends information to each embedded database to inquire about the current transaction status and restrictions of the embedded database. For the embedded database that is in the process of a transaction or has data modification restrictions, compression is not performed. Otherwise, compression is performed;

[0035] The data set classification module is used for the embedded database to be compressed. First, it counts the most recent used transactions of each table in the embedded database. If a table contains data that has never been accessed within more than 100 data-related transactions, the table is marked as an infrequently used table. Then, the K-Means algorithm is used to classify the data in each infrequently used table by primary key or timestamp, and each table is divided into different data sets;

[0036] The data set evaluation module is used to perform feature evaluation at the data level on each data set obtained in the data set classification module, and select the data sets that need to be compressed according to the evaluation results;

[0037] Specifically, the data set evaluation module is as follows:

[0038] Set the read operation transaction metric value of the previous embedded database to increase by 1 for every 10 transaction numbers that have not been read, set the write operation transaction metric value of the previous embedded database to increase by 1 for every 10 transaction numbers that have not been written, record the metric value of the regular structure in the data itself as 1, and the metric value of the irregular structure as 0;

[0039] Set the weight of the read operation to 1, the weight of the write operation to 2, and the structure weight of the data itself to 1. Multiply all metric values by their weights to obtain the total weight value of a data set. Sort the total weight values of all data sets from high to low, and select the data set with the highest total weight value for data compression;

[0040] The Q-learning module is used to use Q-learning as a way to select a compression algorithm. It takes the current CPU, memory, and hard disk usage as the state input of Q-learning, and the selected compression algorithm as the output. It calculates the reward using the throughput improvement amount of the compressed embedded database and the reduction amount of the occupied resources. It obtains the compression algorithm through each iteration, and compresses the data sets obtained by the data set evaluation module;

[0041] The specific steps of the system detection module are as follows:

[0042] Create a daemon process and use it to detect the current CPU, memory, and hard disk usage. If the usage rate of any one of the CPU and memory exceeds 70%, compression will not be performed, and the daemon process will enter the sleep state, waiting to be reawakened for judgment again. If the usage rates of both the CPU and memory do not exceed 70%, compression will be performed;

[0043] The specific steps for creating the daemon process are as follows:

[0044] Use a Python script to create a daemon process that wakes up every 5 minutes. This daemon process will detect the current CPU, memory, and hard disk usage every 5 minutes;

[0045] The specific steps for connecting to the embedded database are as follows:

[0046] Send a connection request to each embedded database through the port number. After the embedded database agrees to the connection, it will connect to the database and ensure the connection of the embedded database through the heartbeat mechanism during subsequent operations;

[0047] The maximum number of data rows in the set in the data set classification module is 1000 rows;

[0048] The specific way of using Q-learning as the method for selecting the compression algorithm in the Q-learning module is as follows:

[0049] When the CPU or memory usage rate reaches more than 50% and less than 70%, select a lightweight compression algorithm;

[0050] When the usage rates of both the CPU and memory are less than 50% and the hard disk usage rate is less than 50%, select a lightweight compression algorithm. If the hard disk usage rate is not less than 50%, select a heavyweight compression algorithm;

[0051] When the hard disk usage rate reaches more than 95%, at this time, the CPU and memory usage rates will be ignored, and the heavyweight algorithm will be directly used to quickly compress the data on the hard disk;

[0052] The lightweight compression algorithms include RLE or delta-of-delta, and the heavyweight compression algorithm is the Gzip algorithm;

[0053] The specific steps for compression in the Q-learning module are as follows:

[0054] When performing compression, first extract the data set to be compressed from the corresponding table in the embedded database, store the extracted data set in a separate file, use the compression algorithm selected by the Q-learning module to compress this file, obtain the compressed file, and finally store the compressed file in an additional folder. This folder stores all the compressed data files;

[0055] The Q-learning module further includes the following steps:

[0056] Record all the compressed data files saved in the folder, and use a B+ tree to index the data files. At the same time, monitor the transactions of the embedded database in real time. Once there is a transaction involving the compressed data files, find the compressed data files through the index, decompress the data files, and then merge the decompressed data back into the table of the embedded database. Finally, update the read or write transaction records of the corresponding data.

[0057] The beneficial effects of the present invention are:

[0058] This application can classify and identify different scenarios and system conditions, and then automatically select the required compression algorithm and automatically decompress when needed. It can adapt to a variety of usage environments, improve the compression speed of the embedded database for different environments, and can intelligently judge the compression timing, such as in device environments like Internet of Things devices, mobile phones, personal computers, etc., and make adaptive adjustments according to the resource quantities in these environments to achieve intelligent compression and decompression of the embedded database, ultimately achieving the purpose of saving storage resource space, improving resource utilization rate, and accelerating network transmission speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 For the process of this application Figure 1 ;

[0060] Figure 2 For the process of this application Figure 2 ;

[0061] Figure 3 For the process of this application Figure 3 . DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] It should be specifically noted that, without conflict, the various embodiments disclosed in this application can be combined with each other.

[0063] DETAILED DESCRIPTION OF THE EMBODIMENT 1: Refer to Figure 1 This embodiment specifically describes the intelligent data compression method for an embedded database, which includes the following steps:

[0064] Step 1: Detect the current usage of the CPU, memory, and hard disk. If the usage rate of any one of the CPU and memory exceeds 70%, no compression is performed. If the usage rates of both the CPU and memory do not exceed 70%, compression is performed, and step 2 is executed;

[0065] Step 2: Determine whether there is an embedded database connection currently. If there is no embedded database connection currently, connect to the embedded database. If there is an embedded database connection currently, send information to each embedded database to inquire about the current transaction status and restrictions of the embedded database. For an embedded database that is in the middle of a transaction or has data modification restrictions, do not perform compression. Otherwise, perform compression and execute Step 3;

[0066] Step 3: For the embedded database that undergoes compression, first count the most recent used transactions of each table in the embedded database. If there is data in the table that has not been accessed within more than 100 data-related transactions, mark the table as an infrequently used table. Then, use the K-Means algorithm to classify the data in each infrequently used table by primary key or timestamp, and divide each table into different data sets;

[0067] Step 4: Perform feature evaluation at the data level on each data set obtained in Step 3, and select the data sets that need to be compressed according to the evaluation results;

[0068] The specific content of Step 4 is as follows:

[0069] Set the read operation transaction metric value of the previous embedded database to increase by 1 for every 10 unread transactions, set the write operation transaction metric value of the previous embedded database to increase by 1 for every 10 unwritten transactions, record the metric value of the regular structure in the data itself as 1 and the metric value of the irregular structure as 0, set the weight of the read operation to 1, the weight of the write operation to 2, and the weight of the data structure itself to 1. Multiply all metric values by their weights to obtain the total weight value of a data set. Sort the total weight values of all data sets from high to low, and select the data set with the highest total weight value for data compression; as Figure 1 shown.

[0070] Step 5: Use Q-learning as a way to select the compression algorithm. Take the current CPU, memory, and hard disk usage as the state input of Q-learning, and take the selected compression algorithm as the output. Calculate the reward using the throughput improvement amount of the compressed embedded database and the reduction amount of the occupied resources, and obtain the compression algorithm through each iteration; as Figure 2 shown.

[0071] Step 6: Compress the data sets obtained in Step 4 using the compression algorithm selected in Step 5. The compression steps are as Figure 3 shown.

[0072] Specific Embodiment 2: This embodiment is a further explanation of Specific Embodiment 1. The difference between this embodiment and Specific Embodiment 1 is that the specific steps of Step 1 are as follows:

[0073] Create a daemon process and use it to detect the current CPU, memory, and hard disk usage. If the usage rate of any one of the CPU and memory exceeds 70%, compression will not be performed, and the daemon process will enter the sleep state, waiting to be woken up again for judgment. If the usage rates of both the CPU and memory do not exceed 70%, compression will be performed, and step two will be executed.

[0074] Specific Embodiment Three: This embodiment is a further description of Specific Embodiment Two. The difference between this embodiment and Specific Embodiment Two is that the specific steps for creating the daemon process are as follows:

[0075] Use a Python script to create a daemon process that wakes up every 5 minutes. This daemon process will detect the current CPU, memory, and hard disk usage every 5 minutes.

[0076] Specific Embodiment Four: This embodiment is a further description of Specific Embodiment Three. The difference between this embodiment and Specific Embodiment Three is that the specific steps for connecting to the embedded database are as follows:

[0077] Send a connection request to each embedded database through the port number. After the embedded database agrees to the connection, it will be connected to the database, and the connection of the embedded database will be ensured through the heartbeat mechanism during subsequent operation.

[0078] Specific Embodiment Five: This embodiment is a further description of Specific Embodiment Four. The difference between this embodiment and Specific Embodiment Four is that the maximum number of data rows in the set in step three is 1000 rows.

[0079] Specific Embodiment Six: This embodiment is a further description of Specific Embodiment Five. The difference between this embodiment and Specific Embodiment Five is that the specific method of using Q-learning as the way to select the compression algorithm in step five is as follows:

[0080] When the CPU or memory usage rate reaches more than 50% and less than 70%, select a lightweight compression algorithm;

[0081] When the usage rates of both the CPU and memory are less than 50% and the hard disk usage rate is less than 50%, select a lightweight compression algorithm. If the hard disk usage rate is not less than 50%, select a heavyweight compression algorithm;

[0082] When the hard disk usage rate reaches more than 95%, at this time, the CPU and memory usage rates will be ignored, and the heavyweight algorithm will be directly used to quickly compress the data on the hard disk.

[0083] Embodiment 7: This embodiment is a further description of Embodiment 6. The difference between this embodiment and Embodiment 6 is that the lightweight compression algorithm includes RLE or delta-of-delta, and the heavyweight compression algorithm is the Gzip algorithm.

[0084] Embodiment 8: This embodiment is a further description of Embodiment 7. The difference between this embodiment and Embodiment 7 is that the specific steps of compression in Step 6 are as follows:

[0085] When performing compression, first extract the data set to be compressed from the corresponding table in the embedded database, store the extracted data set in a separate file, use the compression algorithm selected in Step 5 to compress this file to obtain the compressed completed file, and finally store the compressed completed file in an additional folder, which stores all the compressed data files.

[0086] Embodiment 9: This embodiment is a further description of Embodiment 8. The difference between this embodiment and Embodiment 8 is that the method further includes Step 7, and the specific steps of Step 7 are as follows:

[0087] Record all the compressed data files saved in the folder, and establish an index for the data files using a B+ tree. At the same time, monitor the transactions of the embedded database in real time. Once there is a transaction involving the compressed data files, find the compressed data files through the index, decompress the data files, then merge the decompressed data back into the table of the embedded database, and finally update the read or write transaction records of the corresponding data.

[0088] Embodiment 10: The data intelligent compression system for an embedded database described in this embodiment includes a system detection module, a connection judgment module, a data set classification module, a data set evaluation module, and a Q-learning module;

[0089] The system detection module is used to detect the current usage of the CPU, memory, and hard disk. If the usage rate of any one of the CPU and memory exceeds 70%, compression is not performed. If the usage rates of both the CPU and memory do not exceed 70%, compression is performed;

[0090] The connection judgment module is used to judge whether there is an embedded database connection currently. If there is no embedded database connection currently, connect to the embedded database. If there is an embedded database connection currently, send information to each embedded database to inquire about the current transaction status and restrictions of the embedded database. For an embedded database that is in the middle of a transaction or has data modification restrictions, compression is not performed. Otherwise, compression is performed;

[0091] The data set classification module is used for the embedded database to be compressed. First, it counts the most recent used transactions of each table in the embedded database. If there is data in the table that has not been accessed within more than 100 data-related transactions, the table is marked as an infrequently used table. Then, the K-Means algorithm is used to classify the data in each infrequently used table by primary key or timestamp, and each table is divided into different data sets;

[0092] The data set evaluation module is used to evaluate the characteristics of each data set obtained in the data set classification module at the data level, and select the data sets that need to be compressed according to the evaluation results;

[0093] Specifically, the data set evaluation module is as follows:

[0094] Set the read operation transaction metric value of the previous embedded database to increase by 1 for every 10 unread transactions, set the write operation transaction metric value of the previous embedded database to increase by 1 for every 10 unwritten transactions, and record the metric value of the regular structure in the data itself as 1 and the metric value of the irregular structure as 0.

[0095] Set the weight of the read operation to 1, the weight of the write operation to 2, and the structure weight of the data itself to 1. Multiply all metric values by their weights to obtain the total weight value of a data set. Sort the total weight values of all data sets from high to low, and select the data set with the highest total weight value for data compression;

[0096] The Q-learning module is used to use Q-learning as a way to select a compression algorithm. The current CPU, memory, and hard disk usage are used as the state input of Q-learning, and the selected compression algorithm is used as the output. The reward is calculated using the throughput improvement amount of the compressed embedded database and the reduction amount of the occupied resources. The compression algorithm is obtained through each iteration, and the data sets obtained by the data set evaluation module are compressed;

[0097] The specific steps of the system detection module are as follows:

[0098] Create a daemon process, and use the daemon process to detect the current CPU, memory, and hard disk usage. If the usage rate of any one of the CPU and memory exceeds 70%, no compression is performed, the daemon process enters the sleep state, waits to be awakened again for judgment, and if the usage rates of both the CPU and memory do not exceed 70%, compression is performed;

[0099] The specific steps of creating the daemon process are as follows:

[0100] Create a daemon process that wakes up every 5 minutes using a Python script. This daemon process checks the current CPU, memory, and hard disk usage every 5 minutes;

[0101] The specific steps for connecting to the embedded database are as follows:

[0102] Send a connection request to each embedded database through the port number. After the embedded database agrees to the connection, it will connect to the database and ensure the connection of the embedded database through the heartbeat mechanism during subsequent operations;

[0103] The maximum number of data rows in the collection in the data collection classification module is 1000 rows;

[0104] The specific way of using Q-learning in the Q-learning module as a method for selecting the compression algorithm is as follows:

[0105] When the CPU or memory usage rate reaches more than 50% and less than 70%, select a lightweight compression algorithm;

[0106] When both the CPU and memory usage rates are less than 50% and the hard disk usage rate is less than 50%, select a lightweight compression algorithm. If the hard disk usage rate is not less than 50%, select a heavyweight compression algorithm;

[0107] When the hard disk usage rate reaches more than 95%, at this time, ignore the CPU and memory usage rates and directly use the heavyweight algorithm to quickly compress the data on the hard disk;

[0108] The lightweight compression algorithms include RLE or delta-of-delta, and the heavyweight compression algorithm is the Gzip algorithm;

[0109] The specific steps for compression in the Q-learning module are as follows:

[0110] When performing compression, first extract the data set to be compressed from the corresponding table in the embedded database, store the extracted data set in a separate file, use the compression algorithm selected by the Q-learning module to compress this file to obtain the compressed file, and finally store the compressed file in an additional folder, which stores all the compressed data files;

[0111] The Q-learning module also includes the following steps:

[0112] Record all the compressed data files saved in the folder, and use a B+ tree to build an index for the data files. At the same time, monitor the transactions of the embedded database in real time. Once there is a transaction involving the compressed data files, find the compressed data files through the index, decompress the data files, and then merge the decompressed data back into the table of the embedded database. Finally, update the read or write transaction records of the corresponding data.

[0113] Embodiment: An intelligent data compression method for an embedded database, which includes the following steps:

[0114] S1. Use a Python script to create a daemon process that wakes up every 5 minutes. This daemon process will detect the CPU, memory, and hard disk usage of the current system every 5 minutes. If the current CPU and memory usage rates exceed 70%, no compression will be performed. At this time, it means that there are other programs running in the system, and the system resources are being occupied and there is no power to carry out additional data compression plans. If the CPU and memory usage rates do not exceed 70%, and the current hard disk usage rate exceeds 70%, it means that the current system faces less pressure from other programs, and the hard disk usage rate is relatively high. The necessity and benefit of compressing data are relatively high, and data compression will start. If it is currently selected not to perform compression, the daemon process will directly enter the sleep state and wait to be woken up again after 5 minutes to make another judgment;

[0115] S2. If it is determined in S1 that data compression should be performed, that is, the CPU and memory usage rates do not exceed 70%, which means that the current system level can carry out the data compression plan, then the second step should be to check whether the databases existing in the current system can perform data compression. As an embedded database that runs embedded in a program, there may be multiple database instances running in the same system. If there is no connection to the embedded database currently, a connection request will be sent to each embedded database through a specific port number. After the embedded database agrees to the connection, it will be connected to the database, and the database connection will be ensured through a heartbeat mechanism during subsequent operations. If there is a database connection currently, send information to each database to inquire about the current transaction status and restrictions of the database. When the database is performing a transaction, compression cannot be performed; for a database with data modification restrictions, compression cannot be performed. For a database determined to be non-compressible, the data in the database will be skipped during compression. Otherwise, the data in the database will be selected for compression;

[0116] S3. If it is determined in S2 that the current database is suitable for data compression, first mark the data that has not been accessed in the last 100 data-related transactions in the database. The specific marking method is to process each table as a unit. For each table, count the most recent usage transactions. If there are more than 100 transactions, mark these tables as infrequently used tables. Then, classify the data in each table using the primary key or timestamp. Data with a relatively small difference in primary key values and similar timestamps are grouped into a set using the K-Means algorithm. Each table is divided into different sets, and the maximum number of data rows in each set is 1000. This set is called a data set;

[0117] S4. Evaluate the characteristics at the data level for each data set in each table obtained in S3, specifically including: the number of transactions since the last read operation transaction of the database. The higher this value, the lower the query frequency of all the data in the set by the database, and the less likely it is to be accessed again and decompressed after compression, making it more suitable for compression. The metric value is set to increase by 1 for every 10 transactions without a read operation; the number of transactions since the last write operation transaction. The higher this value, the lower the write operation frequency of the current database for this data. Since changes to a part of the compressed data will cause changes to the entire compressed file and require re-compression, the weight of the write operation is higher than that of the read operation. The metric value is set to increase by 1 for every 10 transactions without a write operation; the structure of the data itself. If the data in the set is stored regularly, such as time-series data in a time-series database, where the length of the timestamp, fields, and labels in each column of each data is similar and changes little, the structure is regular, and the metric value is recorded as 1. During compression, this can achieve the maximum compression ratio by the compression algorithm, reduce more hardware space usage, and is suitable for compressing such data. If the structure is irregular, the metric value is recorded as 0. Among them, the weight of the read operation is 1, the weight of the write operation is 2, and the weight of the structure itself is 1. Multiply these metric values by their weights to obtain the total weight of a data set. Sort these sets in descending order according to the above total weights, and select the current largest and most suitable data set for data compression. Only compress this data set to prevent the compression operation from being overly extensive, which may have a too large impact on the system and is more likely to have data accessed in the future, causing the compressed data to be decompressed. Therefore, only one set is selected for compression;

[0118] S5. The algorithm obtains the CPU, memory, and hard disk utilization rates of the current system, and selects the required data compression method based on these states. Q-learning is used in the algorithm as a way to select the compression algorithm. The state of the system is used as the state input of Q-learning, and the selection of the compression algorithm is used as the action. The reward is calculated using the throughput increase of the compressed database and the reduction of the occupied resources. The most suitable compression algorithm for the current situation is obtained through each iteration. When the CPU and memory utilization rates are relatively high, above 50% and below 70%, it means that there are other processes running in the system currently. At this time, the algorithm should select lightweight compression algorithms such as RLE or delta-of-delta, and the specific algorithm is output by Q-learning. If the CPU and memory utilization rates of the system are relatively low, below 50% at this time, the Q-learning algorithm is more inclined to select heavyweight compression algorithms such as the Gzip algorithm. When the hard disk utilization rate is relatively low, such as below 50%, the Q-learning algorithm is more inclined to use lightweight compression algorithms that save time but have a lower compression ratio. When the hard disk utilization rate is relatively high, such as above 50%, it is more inclined to use heavyweight compression algorithms with a high compression ratio and low efficiency. In addition, when the hard disk utilization rate is very high, such as above 95%, due to the penalty for high hard disk utilization rate in the reward function, the CPU and memory utilization rates of the system will be ignored at this time, and the heavyweight algorithm will be directly used to quickly compress the data on the hard disk to prevent the hard disk utilization rate from reaching 100% and losing data. The alternative compression algorithms include lightweight compression algorithms such as RLE and delta-of-delta, and heavyweight general algorithms such as the Gzip algorithm;

[0119] S6. Use the compression algorithm output by the algorithm in S5 to compress the data set to be compressed in S4. When compressing, first extract the data set to be compressed from the database table, store it in a separate file, use the compression algorithm output in S5 to compress this file and obtain the compressed completed file, and finally store the compressed completed file in an additional folder, which stores all the compressed data files;

[0120] S7. Record the compressed data files obtained after compression in S6, and build indexes for these data files using a B+ tree. At the same time, monitor the transactions of the database in real time. Once there is a transaction involving the compressed data, find the compressed data files from the index file, decompress these data in a timely manner, and re-merge these data into the database table to reduce the impact on these transactions. Finally, update the read or write transaction records of these data.

[0121] It should be noted that the specific implementation manners are only explanations and illustrations of the technical solutions of the present invention, and the scope of the patent protection cannot be limited thereby. Any partial changes made only in accordance with the claims and the description of the present invention shall still fall within the protection scope of the present invention.

Claims

1. An intelligent data compression method for embedded databases, characterized in that it includes the following steps: Step 1: Detect the current usage of CPU, memory, and hard disk. If the usage rate of any one of the CPU and memory exceeds 70%, no compression is performed. If the usage rates of both the CPU and memory do not exceed 70%, compression is performed, and Step 2 is executed; Step 2: Determine whether there is an embedded database connection currently. If there is no embedded database connection currently, connect to the embedded database. If there is an embedded database connection currently, send information to each embedded database to inquire about the current transaction status and restrictions of the embedded database. For an embedded database that is in the middle of a transaction or has data modification restrictions, no compression is performed. Otherwise, compression is performed, and Step 3 is executed; Step 3: For the embedded database to be compressed, first count the most recent usage transactions of each table in the embedded database. If there is data in the table that has not been accessed in more than 100 data-related transactions, mark the table as an infrequently used table. Then, use the K-Means algorithm to classify the data in each infrequently used table by primary key or timestamp, and divide each table into different data sets; Step 4: Perform feature evaluation at the data level for each data set obtained in Step 3, and select the data sets that need to be compressed according to the evaluation results; The specific content of Step 4 is as follows: Set the read operation transaction metric value of the previous embedded database to increase by 1 for every 10 unread transactions, set the write operation transaction metric value of the previous embedded database to increase by 1 for every 10 unwritten transactions, and record the metric value of the regular structure in the data itself as 1 and the metric value of the irregular structure as 0. Set the weight of the read operation to 1, the weight of the write operation to 2, and the weight of the data structure itself to 1. Multiply all metric values by their weights to obtain the total weight value of a data set. Sort the total weight values of all data sets from high to low, and select the data set with the highest total weight value for data compression; Step 5: Use Q-learning as a way to select the compression algorithm. Take the current usage of CPU, memory, and hard disk as the state input of Q-learning, and take the selected compression algorithm as the output. Calculate the reward using the throughput improvement amount of the compressed embedded database and the reduction amount of occupied resources, and obtain the compression algorithm through each iteration; Step 6: Compress the data sets obtained in Step 4 using the compression algorithm selected in Step 5; The specific content of using Q-learning as a way to select the compression algorithm in Step 5 is as follows: When the CPU or memory usage rate reaches more than 50% and less than 70%, select a lightweight compression algorithm; When the usage rates of both the CPU and memory are less than 50%, and the hard disk usage rate is less than 50%, select a lightweight compression algorithm. If the hard disk usage rate is not less than 50%, select a heavyweight compression algorithm; When the hard disk usage rate reaches more than 95%, at this time, ignore the CPU and memory usage rates, and directly use the heavyweight algorithm to quickly compress the data on the hard disk.

2. The data intelligent compression method for an embedded database according to claim 1, characterized in that the specific steps of step one are as follows: Create a daemon process, and use the daemon process to detect the current CPU, memory, and hard disk usage. If the usage rate of any one of the CPU and memory exceeds 70%, compression is not performed, the daemon process enters the sleep state, waits to be awakened again for judgment, and if the usage rates of both the CPU and memory do not exceed 70%, compression is performed, and step two is executed.

3. The data intelligent compression method for an embedded database according to claim 2, characterized in that the specific steps of creating the daemon process are as follows: Use a Python script to create a daemon process that wakes up every 5 minutes. This daemon process will detect the current CPU, memory, and hard disk usage every 5 minutes.

4. The data intelligent compression method for an embedded database according to claim 3, characterized in that the specific steps of connecting to the embedded database are as follows: Send a connection request to each embedded database through the port number. After the embedded database agrees to the connection, it will connect to the database, and in subsequent operations, ensure the connection of the embedded database through the heartbeat mechanism.

5. The data intelligent compression method for an embedded database according to claim 4, characterized in that the maximum number of data rows in the set in step three is 1000 rows.

6. The data intelligent compression method for an embedded database according to claim 5, characterized in that the lightweight compression algorithm includes RLE or delta-of-delta, and the heavyweight compression algorithm is the Gzip algorithm.

7. The data intelligent compression method for an embedded database according to claim 6, characterized in that the specific steps of compression in step six are as follows: When performing compression, first extract the data set to be compressed from the corresponding table in the embedded database, store the extracted data set in a separate file, use the compression algorithm selected in step five to compress this file, obtain the compressed completed file, and finally store the compressed completed file in an additional folder, and all compressed data files are saved in this folder.

8. The data intelligent compression method for an embedded database according to claim 7, characterized in that the method further includes step seven, and the specific steps of step seven are as follows: Record all the compressed data files saved in the folder, and use a B+ tree to establish an index for the data files. At the same time, monitor the transactions of the embedded database in real time. Once there is a transaction involving the compressed data files, find the compressed data files through the index, decompress the data files, then merge the decompressed data back into the table of the embedded database, and finally update the read or write transaction records of the corresponding data.

9. A data intelligent compression system for an embedded database, which is applied to the compression method described in any one of claims 1 to 8, characterized in that The system includes a system detection module, a connection judgment module, a data set classification module, a data set evaluation module, and a Q-learning module; The system detection module is used to detect the current CPU, memory, and hard disk usage. If the usage rate of any one of the CPU and memory exceeds 70%, compression is not performed. If the usage rates of both the CPU and memory do not exceed 70%, compression is performed; The connection judgment module is used to judge whether there is an embedded database connection currently. If there is no embedded database connection currently, the embedded database is connected. If there is an embedded database connection currently, information is sent to each embedded database to inquire about the current transaction status and restrictions of the embedded database. For the embedded database that is in the middle of a transaction or has data modification restrictions, compression is not performed. Otherwise, compression is performed; The data set classification module is used for the embedded database to be compressed. First, the most recent usage transactions of each table in the embedded database are counted. If there is data that has never been accessed in more than 100 data-related transactions in the table, the table is marked as an infrequently used table. Then, the K-Means algorithm is used for the data in each infrequently used table to classify the primary key or timestamp, and each table is divided into different data sets; The data set evaluation module is used to perform feature evaluation at the data level for each data set obtained in the data set classification module, and select the data sets that need to be compressed according to the evaluation results; The Q-learning module is used to use Q-learning as a way to select the compression algorithm. The current CPU, memory, and hard disk usage are used as the state input of Q-learning, and the selected compression algorithm is used as the output. The reward is calculated using the throughput improvement amount of the compressed embedded database and the reduction amount of the occupied resources. The compression algorithm is obtained through each iteration, and the data sets obtained by the data set evaluation module are compressed.

Citation Information

Patent Citations

  • Storage space optimization method and system for database files

    CN114564457A

  • Method for data compression transmission considering bandwidth in hadoop cluster, recording medium and device for performing the method

    KR102195239B1