Database time sequence data storage method and related product

By performing segmented difference encoding on time series data and initial compression of different compression algorithms, the problem of poor compression effect of differential encoding algorithms when the timing data fluctuates greatly, achieving more efficient data compression and transmission.

CN120123342APending Publication Date: 2025-06-10CETC JINCANG (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510193219.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When the existing differential encoding algorithms have large fluctuations or jumps in the time series data value, the compression effect is poor, making it difficult to effectively reduce the storage space of the time series data and improve the transmission efficiency.

Method used

By analyzing the numerical fluctuations of the time series data, it is segmented into a first data segment with a large numerical fluctuations and a second data segment with a small numerical fluctuations, and the difference encoding and initial compression of different preset compression algorithms are performed respectively.

Benefits of technology

It improves the overall compression effect and compression efficiency of timing data, reduces the storage space of the database and improves the data transmission efficiency, thereby improving the database's storage and query performance of timing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123342A_ABST
    Figure CN120123342A_ABST
Patent Text Reader

Abstract

The invention provides a database time series data storage method and a related product. The storage method comprises the following steps: analyzing numerical fluctuation of acquired time series data; segmenting the time series data according to the magnitude of the numerical value fluctuation to obtain at least one first data segment and at least one second data segment, the numerical value fluctuation of the first data segment being greater than the numerical value fluctuation of the second data segment; performing difference value coding on the first data segment to obtain a first difference value sequence, and performing difference value coding on the second data segment to obtain a second difference value sequence; performing primary compression on the first difference value sequence and the second difference value sequence through different preset compression algorithms; and storing the data obtained by primary compression. According to the method, the compression effect and the compression efficiency of the time series data are integrally improved, the storage space of the database is reduced, the data transmission efficiency of the database is improved, and the storage and query performance of the database on the time series data is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of databases, and in particular, to a method for storing time-series data in a database, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the explosive growth of digital information, databases need to store and process a large amount of time-series data (hereinafter referred to as time-series data). In related technologies, the differential encoding algorithm (Delta Encoding) is used to perform differential encoding on time-series data and then store it in a database. This solution can reduce the storage space of time-series data and improve the transmission efficiency of time-series data. However, when there are large fluctuations or jumps in the values of time-series data, the compression effect of the differential encoding algorithm is poor, and it is difficult to effectively reduce the storage space of time-series data and improve the transmission efficiency of time-series data. Summary of the Invention

[0003] An object of the present invention is to provide a method for storing time-series data in a database, a computer-readable storage medium, and a computer program product, so as to improve the compression effect on time-series data, thereby reducing the storage space of time-series data and improving the transmission efficiency of time-series data.

[0004] A further object of the present invention is to reduce or avoid the influence of outliers in time-series data on compression based on the differential encoding algorithm, and further improve the compression effect of time-series data.

[0005] Specifically, according to one aspect of the present invention, the present invention provides a method for storing time-series data in a database, including:

[0006] Analyze the numerical fluctuations of the acquired time-series data;

[0007] Segment the time-series data according to the magnitude of the numerical fluctuations to obtain at least one first data segment and at least one second data segment, wherein the numerical fluctuations of the first data segment are greater than those of the second data segment;

[0008] Perform differential encoding on the first data segment to obtain a first difference sequence, and perform differential encoding on the second data segment to obtain a second difference sequence;

[0009] Perform primary compression on the first difference sequence and the second difference sequence respectively through different preset compression algorithms;

[0010] Store the data obtained by primary compression.

[0011] Optionally, the step of performing primary compression on the second difference sequence through a preset compression algorithm includes:

[0012] Determine whether there are multiple sections with consecutive identical values in the second difference sequence;

[0013] If so, compress the second difference sequence through a run-length encoding algorithm;

[0014] If not, compress the second difference sequence through a variable-length encoding algorithm; the variable-length encoding algorithm includes Elias gamma code algorithm or Huffman coding algorithm.

[0015] Optionally, the step of initially compressing the first difference sequence through a preset compression algorithm includes:

[0016] Compress the first difference sequence through a two's complement encoding algorithm.

[0017] Optionally, before initially compressing the first difference sequence and the second difference sequence respectively through different preset compression algorithms, the storage method further includes:

[0018] Based on a preset statistical method, determine whether there are outliers in the first data segment and the second data segment;

[0019] If there are outliers, determine a repair value according to the value of the time-series data adjacent to the outlier and a preset interpolation algorithm, and replace the outlier with the repair value.

[0020] Optionally, the step of segmenting the time-series data according to the magnitude of the numerical fluctuation to obtain at least one first data segment and at least one second data segment includes:

[0021] Obtain a preset numerical fluctuation threshold;

[0022] Divide the time-series data into at least one first data segment and at least one second data segment according to the preset numerical fluctuation threshold and the numerical fluctuation of the values at each place in the time-series data.

[0023] Optionally, the step of performing difference encoding on the first data segment to obtain a first difference sequence and performing difference encoding on the second data segment to obtain a second difference sequence includes:

[0024] Perform difference encoding on the first data segment with a first encoding bit number to obtain a first difference sequence; and

[0025] Perform difference encoding on the second data segment with a second encoding bit number to obtain a second difference sequence; wherein, the first encoding bit number is greater than the second encoding bit number.

[0026] Optionally, storing the data obtained by initial compression includes:

[0027] Determine the compression level corresponding to the time-series data according to the timeliness and / or expected access frequency of the time-series data. The compression level includes a first compression level and a second compression level, and the first compression level is greater than the second compression level;

[0028] Re-compress the time-series data after the initial compression according to the compression level;

[0029] Store the time-series data compressed according to the first compression level into a first storage medium; and

[0030] Store the time-series data compressed according to the second compression level into a second storage medium; wherein the read / write speed of the first storage medium is less than that of the second storage medium.

[0031] Optionally, the determining the compression level corresponding to the time-series data according to the timeliness and / or expected access frequency of the time-series data includes:

[0032] Judge whether the time-series data belongs to streaming data or historical data according to the timeliness;

[0033] If it belongs to the streaming data, set the time-series data to the first compression level;

[0034] If it belongs to the historical data, set the time-series data to the second compression level; and

[0035] Judge whether the time-series data belongs to high-frequency access data or low-frequency access data according to the expected access frequency;

[0036] If it belongs to the high-frequency access data, set the time-series data to the second compression level;

[0037] If it belongs to the low-frequency access data, set the time-series data to the first compression level.

[0038] According to another aspect of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the storage method of the database time-series data described above are implemented.

[0039] According to still another aspect of the present invention, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the storage method of the database time-series data described above are implemented.

[0040] The storage method for database time-series data of the present invention divides the time-series data into a first data segment with relatively large numerical fluctuations and a second data segment with relatively small numerical fluctuations, and performs primary compression using different preset compression algorithms according to the characteristics of the first difference sequence and the second difference sequence after difference encoding. Overall, it improves the compression effect and efficiency of the time-series data, reduces the database storage space, improves the database data transmission efficiency, and further improves the storage and query performance of the database for time-series data.

[0041] From the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more clear about the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Some specific embodiments of the present invention will be described in detail hereinafter with reference to the accompanying drawings in an exemplary but not restrictive manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0043] Figure 1 is a flowchart of the storage method according to an embodiment of the present invention;

[0044] Figure 2 is a flowchart of segmenting time-series data in the storage method according to an embodiment of the present invention;

[0045] Figure 3 is a flowchart of performing difference encoding on time-series data in the storage method according to an embodiment of the present invention;

[0046] Figure 4 is a flowchart of performing primary compression on time-series data in the storage method according to an embodiment of the present invention;

[0047] Figure 5 is a flowchart of repairing outliers in the storage method according to an embodiment of the present invention;

[0048] Figure 6 is a flowchart of storing time-series data in the storage method according to an embodiment of the present invention;

[0049] Figure 7 is a flowchart of determining the compression level in the storage method according to an embodiment of the present invention;

[0050] Figure 8 is a flowchart of determining the compression level in the storage method according to another embodiment of the present invention;

[0051] Figure 9Schematic diagram of a computer-readable storage medium according to an embodiment of the present invention; and

[0052] Figure 10 Schematic diagram of a computer program product according to an embodiment of the present invention. Detailed implementation manners

[0053] The purpose of the method for storing database time-series data in this embodiment is to improve the compression effect on time-series data, thereby reducing the storage space of time-series data and improving the transmission efficiency of time-series data.

[0054] Figure 1 Flow schematic diagram of a storage method according to an embodiment of the present invention, which generally may include:

[0055] S100, analyzing the numerical fluctuations of the acquired time-series data;

[0056] S200, segmenting the time-series data according to the magnitude of the numerical fluctuations to obtain at least one first data segment and at least one second data segment, where the numerical fluctuations of the first data segment are greater than those of the second data segment;

[0057] S300, performing differential encoding on the first data segment to obtain a first difference sequence, and performing differential encoding on the second data segment to obtain a second difference sequence;

[0058] S400, respectively performing primary compression on the first difference sequence and the second difference sequence through different preset compression algorithms;

[0059] S500, storing the data obtained by the primary compression.

[0060] Time-series data, also known as time-sequence data, refers to a data sequence arranged in chronological order, where each data point in the data sequence is associated with a specific timestamp. Time-series data widely exists in many fields, such as stock price fluctuations in the financial market, data records of sensors (such as temperature, humidity, pressure, etc.), weather data, readings of Internet of Things devices, etc. Time-series data has the following main characteristics: time dependence, each data point in time-series data is related not only to its value but also to the data points before and after it; autocorrelation, the values at certain time points may affect the future values; periodicity or trend, time-series data usually exhibits periodicity (such as seasonal changes) and long-term trends (such as economic growth trends); having noise and outliers, there are often some noises (also known as outliers) in time-series data, and these data may require special processing or filtering.

[0061] The differential encoding algorithm is a commonly used method for processing time-series data. It is used to record the differences between adjacent values in a data sequence rather than the values of each data point. Its core idea is to compress data by recording the amount of data change (i.e., "Delta"). Specifically, the differential encoding algorithm will first record the value of the first data point in the data sequence, and then sequentially record the difference between each subsequent data point and the previous data point. Generally, the values of adjacent or nearby data points in time-series data tend to change less, that is, the differences are smaller. Therefore, the differences can usually be represented by fewer bits (i.e., bits / Bit), thus reducing the storage space of the data. When the values in the time-series data change little and the change pattern is stable, the compression effect of the differential encoding algorithm is particularly significant, which will greatly reduce the storage space of the time-series data and improve the transmission efficiency of the time-series data, thereby reducing the cost of the database on storage devices and the cost on network transmission devices. However, when there are large fluctuations or jumps in the values of the time-series data, the differences between adjacent data points may become very large, and a larger number of bits are required to represent the differences, resulting in a significant decline in the compression effect.

[0062] In this embodiment, the time-series data can be analyzed by statistical methods or the like to obtain the numerical fluctuations of the values at various locations in the time-series data. The statistical methods can be preset thresholds, means, standard deviations, data trends, etc. The numerical fluctuations can characterize the fluctuation magnitude of the values of a certain adjacent data point.

[0063] After obtaining the numerical fluctuations of the values at various locations in the time-series data, the time-series data can be divided into multiple data segments according to the numerical fluctuations, where the first data segment has larger numerical fluctuations and the second data segment has smaller numerical fluctuations. The first data segment can be one or more, and the second data segment can be one or more. Generally, the first data segment and the second data segment can be arranged alternately to maintain the continuity of the data sequence in each data segment.

[0064] Next, the differential encoding algorithm can be used to perform difference encoding on each first data segment and second data segment to obtain the difference sequences corresponding to each first data segment and second data segment respectively. It should be understood that the absolute values of the differences in the first difference sequence are relatively large, and the absolute values of the differences in the second difference sequence are relatively small.

[0065] In this embodiment, according to the characteristics of the first difference sequence and the second difference sequence respectively, different preset compression algorithms are used to perform primary compression on the first difference sequence and the second difference sequence respectively. That is to say, the compression effect and compression efficiency are further enhanced through hybrid coding technology. Exemplarily, for the second difference sequence, variable-length coding algorithms, run-length encoding algorithms (Run-Length Encoding, abbreviated as RLE), etc. can be used for primary compression to further reduce the storage space. For the first data segment, complementary codes or other efficient binary coding methods can be used to find a balance between compression efficiency and storage space. In this way, the compression effect and compression efficiency of time-series data are improved as a whole, the storage space of the database is reduced, the data transmission efficiency of the database is improved, and thus the storage and query performance of the database for time-series data is improved.

[0066] After the primary compression of the time-series data, the data obtained after the primary compression can be processed and stored according to the preset storage strategy. Exemplarily, the preset storage strategy may include a compression level strategy and a storage medium strategy. The compression level strategy is used to determine the compression level, target compression ratio, etc. corresponding to the time-series data to determine whether multiple compressions of the time-series data are required. The storage medium strategy is used to determine the type of storage medium corresponding to the time-series data to determine whether the time-series data is stored in an efficient storage medium or an inefficient storage medium.

[0067] In some embodiments of the storage method of the present invention, such as Figure 2 shown, the segmentation of the time-series data according to the magnitude of the numerical fluctuation to obtain at least one first data segment and at least one second data segment includes:

[0068] S211, obtaining a preset numerical fluctuation threshold;

[0069] S213, dividing the time-series data into at least one first data segment and at least one second data segment according to the preset numerical fluctuation threshold and the numerical fluctuation of the values at various places in the time-series data.

[0070] The numerical fluctuation can be an absolute value or a relative value. The preset numerical fluctuation threshold can be an absolute value or a relative value corresponding to the numerical fluctuation, which can be set according to actual needs. In this embodiment, the obtained numerical fluctuation can be compared with the preset numerical fluctuation threshold, and the continuous sequence in the time-series data that exceeds the preset numerical fluctuation threshold is divided into the first data segment, and the continuous sequence that does not exceed the preset numerical fluctuation threshold is divided into the second data segment. It should be understood that each first data segment and each second data segment are continuous sequence data segments.

[0071] In some embodiments of the storage method of the present invention, such as Figure 3As shown, obtaining a first difference sequence by performing difference encoding on a first data segment and obtaining a second difference sequence by performing difference encoding on a second data segment includes:

[0072] S311, performing difference encoding on the first data segment with a first encoding bit number to obtain a first difference sequence; and

[0073] S313, performing difference encoding on the second data segment with a second encoding bit number to obtain a second difference sequence; wherein, the first encoding bit number is greater than the second encoding bit number.

[0074] The value of the first data segment fluctuates greatly, and the value of the difference sequence after difference encoding is also large, so a larger encoding bit number is required. The value of the second data segment fluctuates less, and the value of the difference sequence after difference encoding is also small. Using a smaller encoding bit number can save storage space. Exemplarily, the first encoding bit number can be 16 bits, and the second encoding bit number can be 8 bits.

[0075] In some embodiments of the storage method of the present invention, as Figure 4 shown, the steps of initially compressing the second difference sequence by a preset compression algorithm include:

[0076] S411, determining whether there are multiple sections with continuously identical values in the second difference sequence;

[0077] S413, if so, compressing the second difference sequence by a run-length encoding algorithm;

[0078] S415, if not, compressing the second difference sequence by a variable-length encoding algorithm; the variable-length encoding algorithm includes Elias Gamma Code or Huffman Coding.

[0079] The run-length encoding algorithm is a simple and effective data compression algorithm, especially suitable for data with continuously repeated values. It reduces data storage space by recording continuously identical values (referred to as "runs") and the number of repetitions. In this embodiment, the run-length encoding algorithm can compress the continuously repeated values in the second difference sequence into one value and its number of repetitions, which can significantly compress the data. The run-length encoding algorithm also has the advantages of simple algorithm and not relying on a complex coding table, so it is more suitable for initially compressing the second data segment with multiple sections having continuously identical values.

[0080] Variable - length coding can flexibly represent differences, thus saving storage space. Among them, the Elias - gamma coding algorithm is a compression algorithm for lossless data compression, which is particularly suitable for efficiently encoding positive integers. It belongs to a general integer coding technology that can compactly represent integer values and is often used in entropy coding, information retrieval, and data compression. In this embodiment, in the case where the second difference sequence has a section with more monotonically increasing and / or monotonically decreasing values (the extended algorithm based on the Elias - gamma coding algorithm can be used to process negative integers), the Elias - gamma coding algorithm can be selected for primary compression to improve the compression effect and save storage space.

[0081] The Huffman coding algorithm is a compression algorithm for lossless data compression, an optimal prefix - coding method based on a greedy strategy, and is often used for efficiently compressing characters, data blocks, or files. In this embodiment, in the case where the values in the second difference sequence are unevenly distributed and there is a large amount of data in the second difference sequence, the Huffman coding algorithm can be selected for primary compression to improve the compression effect and save storage space.

[0082] In some embodiments of the storage method of the present invention, the step of primarily compressing the first difference sequence by a preset compression algorithm includes:

[0083] Compress the first difference sequence by the two's - complement coding algorithm.

[0084] The two's - complement coding algorithm has wide applicability in data compression. Especially when dealing with continuously changing integer sequences or signed data, it can efficiently represent signed numbers and reduce storage space. In this embodiment, due to large numerical fluctuations, there must be many negative numbers in the first difference sequence, and the absolute values of the values in the first difference sequence are relatively large. By using the two's - complement coding algorithm to primarily compress the first difference sequence, the compression effect can be effectively improved, and a balance can be found between the compression effect and storage space.

[0085] In some embodiments of the storage method of the present invention, as Figure 5 shown, before primarily compressing the first difference sequence and the second difference sequence by different preset compression algorithms respectively, the storage method further includes:

[0086] S611, determine whether there are outliers in the first data segment and the second data segment;

[0087] S613, if there is an outlier, determine a repair value according to the value of the time - series data adjacent to the outlier and a preset interpolation algorithm, and replace the outlier with the repair value.

[0088] Time series data usually reflects the dynamic behavior of a system or process over time, and time series data has sequentiality or trend, and there is usually a dependency relationship between adjacent or nearby data points. Outliers will break the sequentiality or trend of time series data. For example, the value deviates greatly from its adjacent data points, and the degree of deviation exceeds the acceptable range. Outliers are usually caused by errors or mistakes. On the one hand, they cannot correctly reflect the real trend. On the other hand, they will break the coherence of the difference sequence during difference coding, thereby affecting the compression effect.

[0089] In this embodiment, it is possible to determine whether there are outliers in the first data segment and the second data segment through a preset statistical method based on the mean value, standard deviation, data trend, etc. The preset statistical method can detect those outliers that deviate from the normal data pattern. For example, the mean value of the whole or each section of the first data segment and the second data segment can be calculated first, and an outlier threshold is set. If the difference between a certain data point and the adjacent data points exceeds the outlier threshold, it is marked as an outlier. Another example is that the outliers in time series data can be judged by the standard score algorithm (also known as the z-score algorithm).

[0090] For the detected outliers, a preset interpolation algorithm can be used for repair. The preset interpolation algorithm can be linear interpolation, mean interpolation, etc. Predict the value that conforms to the data trend according to the values of the data points before and after the outlier, and use this value as the repair value to replace the outlier. By searching for and repairing outliers in time series data, the quality and compression effect of time series data can be improved.

[0091] In some embodiments of the storage method of the present invention, as Figure 6 shown, storing the data obtained by the initial compression includes:

[0092] S511, determining the compression level corresponding to the time series data according to the timeliness and / or expected access frequency of the time series data. The compression level includes a first compression level and a second compression level, and the first compression level is greater than the second compression level;

[0093] S513, recompressing the time series data after the initial compression according to the compression level;

[0094] S515, storing the time series data compressed by the first compression level into the first storage medium; and

[0095] S517, storing the time series data compressed by the second compression level into the second storage medium; wherein the read / write speed of the first storage medium is less than the read / write speed of the second storage medium.

[0096] Generally speaking, the larger the compression level, the higher the compression ratio of the data, and the more storage space is saved. However, the larger the compression level, the more time it takes for compression and the more time it takes for decompression. In this embodiment, different compression levels and storage media are determined according to the different timeliness and / or expected access frequencies of different time series data. For data with a relatively low expected access frequency, a larger compression level can be set and stored in a storage medium with a relatively low read and write speed (usually with a lower cost) to save storage space, improve data transmission efficiency, and save the cost of storage devices. For data with a relatively high expected access frequency, a smaller compression level can be set and stored in a storage medium with a relatively high read and write speed to increase the speed of reading and decompressing data during access and reduce the user waiting time.

[0097] It should be understood that both the first compression level and the second compression level include initial compression. For example, when the second compression level is 2, only one more compression is required for the time series data after the initial compression.

[0098] In some embodiments of the storage method of the present invention, as Figures 7 - 8 shown, determining the compression level corresponding to the time series data according to the timeliness and / or expected access frequency of the time series data includes:

[0099] S521, judging whether the time series data belongs to streaming data or historical data according to the timeliness;

[0100] S523, if it belongs to streaming data, setting the time series data to the first compression level;

[0101] S525, if it belongs to historical data, setting the time series data to the second compression level. And

[0102] S531, judging whether the time series data belongs to high-frequency access data or low-frequency access data according to the expected access frequency;

[0103] S533, if it belongs to high-frequency access data, setting the time series data to the second compression level;

[0104] S535, if it belongs to low-frequency access data, setting the time series data to the first compression level.

[0105] Streaming data usually belongs to low-frequency access data and can be compressed using a higher compression level to save storage space. Historical data usually belongs to high-frequency access data and can be compressed using a lower compression level to save compression and decompression time. In this embodiment, for recently arrived streaming data, it is not necessary to recompress the entire data set each time. Only the newly arrived time-series data can be differentially encoded and compressed. This approach can significantly reduce the computational overhead and improve the real-time performance of compression. It can also reduce latency. Especially when processing high-frequency data streams, it can quickly respond to the compression requirements of new data, enhancing the real-time performance of the database system. Additionally, for high-frequency access data, the time overhead during the read and decompression processes can be greatly reduced through the optimization of parallel decompression and caching mechanisms. Especially in the scenario of reading large-scale data sets, the decompression speed can be effectively improved.

[0106] The flowcharts provided in this embodiment are not intended to indicate that the operations of the method will be executed in any specific order, or that all operations of the method are included in every case. Additionally, the method may include additional operations. Within the scope of the technical concept provided by the method of this embodiment, additional changes can be made to the above method.

[0107] It should be understood that in some embodiments, each part can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system.

[0108] This embodiment also provides a computer program product 10 and a computer-readable storage medium 20. Figure 10 is a schematic diagram of a computer program product 10 according to an embodiment of the present invention, Figure 9 is a schematic diagram of a computer-readable storage medium 20 according to an embodiment of the present invention. The computer program product 10 includes a computer program 11. When the computer program 11 is executed by a processor 32, it implements the steps of any one of the above storage methods. The computer-readable storage medium 20 stores the above computer program 11. When the computer program 11 is executed by a processor 32, it implements the steps of any one of the above embodiments of the storage method.

[0109] The computer program 11 for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, configuration data of an integrated circuit, or source code or object code written in any combination of one or more programming languages and procedural programming languages. The computer program 11 may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using the Internet through an Internet service provider). In some embodiments, in order to perform aspects of the present invention, an electronic circuit, including for example a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.

[0110] For the description of this embodiment, the computer program product 10 is a related product containing the computer program 11.

[0111] For the description of this embodiment, the computer-readable storage medium 20 is a tangible device capable of retaining and storing the computer program 11, which may be any device that can contain, store, communicate, propagate, or transport the computer program 11 for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable storage medium 20 include the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device, and any suitable combination of the above.

[0112] The computer program product 10 can run on a computer device. The computer device may include a memory, a processor 32, and a computer program 11 stored on the memory and running on the processor 32. The computer device may be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smart phone. In some examples, the computer device may be a cloud computing node. The computer device may be described in the general context of computer system executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, object programs, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. The computer device may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules may be located on local or remote computing system storage media including storage devices.

[0113] The computer device may include a processor 32 adapted to execute stored instructions and a memory that provides temporary storage space for the operation of the instructions during operation. The processor 32 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory may include random access memory (RAM), read-only memory, flash memory, or any other suitable storage system.

[0114] The computer device may further include a network adapter / interface and an input / output (I / O) interface. The I / O interface allows data to be input and output with external devices that can be connected to the computer device. The network adapter / interface may provide communication between the computer device and a network, which is typically shown as a communication network.

[0115] At this point, those skilled in the art should recognize that although multiple exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications that conform to the principles of the present invention can still be directly determined or derived from the content disclosed in the present invention without departing from the spirit and scope of the present invention. Therefore, the scope of the present invention should be understood and construed to cover all such other variations or modifications.

Claims

1. A method for storing time series data in a database, characterized in that: include: Analyze the numerical fluctuations of acquired time series data; Segmenting the time series data according to the magnitude of the value fluctuation to obtain at least one first data segment and at least one second data segment, wherein the value fluctuation of the first data segment is greater than the value fluctuation of the second data segment; Performing difference encoding on the first data segment to obtain a first difference sequence, and performing difference encoding on the second data segment to obtain a second difference sequence; Performing initial compression on the first difference sequence and the second difference sequence respectively by using different preset compression algorithms; The data obtained by the initial compression is stored.

2. The storage method according to claim 1, characterized in that: The step of initially compressing the second difference sequence by using a preset compression algorithm comprises: Determine whether there are multiple segments with consecutive identical values ​​in the second difference sequence; If it exists, compressing the second difference sequence by a run length encoding algorithm; If it does not exist, the second difference sequence is compressed by a variable length coding algorithm; the variable length coding algorithm includes Elijah Gamma coding algorithm or Huffman coding algorithm.

3. The storage method according to claim 1, characterized in that: The step of initially compressing the first difference sequence by using a preset compression algorithm comprises: The first difference sequence is compressed by a binary complement encoding algorithm.

4. The storage method according to claim 1, characterized in that: Before the first difference sequence and the second difference sequence are respectively compressed initially by using different preset compression algorithms, the storage method further includes: Determine whether there are abnormal values ​​in the first data segment and the second data segment; If an abnormal value exists, a repair value is determined according to the value of the time series data adjacent to the abnormal value and a preset interpolation algorithm, and the abnormal value is replaced by the repair value.

5. The storage method according to claim 1, characterized in that: The segmenting of the time series data according to the magnitude of the value fluctuation to obtain at least one first data segment and at least one second data segment includes: Get the preset value fluctuation threshold; According to the preset numerical fluctuation threshold and the numerical fluctuation of the values ​​at various locations in the time series data, the time series data is divided into at least one first data segment and at least one second data segment.

6. The storage method according to claim 1, characterized in that: Performing difference encoding on the first data segment to obtain a first difference sequence, and performing difference encoding on the second data segment to obtain a second difference sequence, comprising: performing difference encoding on the first data segment using a first encoding bit number to obtain a first difference sequence; and The second data segment is difference-encoded using a second coding bit number to obtain a second difference sequence; wherein the first coding bit number is greater than the second coding bit number.

7. The storage method according to claim 1, characterized in that: The storing of the data obtained by the initial compression includes: Determine, according to the timeliness and / or expected access frequency of the time series data, a compression level corresponding to the time series data, the compression level comprising a first compression level and a second compression level, the first compression level being greater than the second compression level; recompressing the time series data after initial compression according to the compression level; storing the time series data compressed according to the first compression level into a first storage medium; and The time series data compressed according to the second compression level is stored in a second storage medium; wherein the read and write speed of the first storage medium is lower than the read and write speed of the second storage medium.

8. The storage method according to claim 7, characterized in that: The step of determining the compression level corresponding to the time series data according to the timeliness and / or expected access frequency of the time series data includes: Determining whether the time series data is streaming data or historical data according to the timeliness; If it belongs to the streaming data, setting the time series data to the first compression level; If it is the historical data, setting the time series data to the second compression level; and Determining whether the time series data is high-frequency access data or low-frequency access data according to the expected access frequency; If it is the high-frequency access data, setting the time series data to the second compression level; If it belongs to the low-frequency access data, the time series data is set to the first compression level.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method for storing database time series data as described in any one of claims 1 to 8 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for storing database time series data as described in any one of claims 1 to 8 are implemented.