Data compression method and device based on abnormal separation, electronic equipment and storage medium

By dividing the data sequence to be compressed into two data sequences of normal value and outlier value, and performing data compression separately, the problem of increasing data compression costs caused by outliers in the prior art is solved, and more efficient data compression is achieved.

CN120128186AActive Publication Date: 2025-06-10TSINGHUA UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510056518.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-10
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In the data compression process, due to the existence of extreme outliers, the time and space cost of data compression increases and the efficiency is not high.

Method used

By dividing the data sequence to be compressed into at least two data sequences according to normal values ​​and outliers, and compressing these data sequences separately, separate compression storage of outliers is achieved.

Benefits of technology

It reduces the time and space cost of data compression, improves data compression efficiency, and effectively solves the problem of increasing compression costs caused by outliers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128186A_ABST
    Figure CN120128186A_ABST
Patent Text Reader

Abstract

The invention provides a data compression method and device based on abnormal separation, electronic equipment and a storage medium, and relates to the technical field of data compression. The method comprises the steps that a to-be-compressed data sequence is divided into at least two data sequences according to a normal value and an abnormal value in the to-be-compressed data sequence, one of the at least two data sequences contains the normal value, and the other data sequences contain the abnormal value; and respectively carrying out data compression on the at least two data sequences. According to the method, the normal value and the abnormal value in the to-be-compressed data sequence are respectively subjected to data compression after being separated, so that the time and space cost of data compression can be reduced, and the data compression efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data compression, and in particular, to a data compression method, device, electronic device and storage medium based on anomaly separation. Background Art

[0002] The Internet of Things and time series analysis are two rapidly developing fields at present, and they play important roles in many applications. With the popularization of Internet of Things devices, the types and quantities of time series data have experienced explosive growth, and the efficient compression and storage of time series data have become increasingly important. In the work of compressing data values by traditional compression algorithms, bit padding is a classic task. Many classic algorithms use bit padding as a component in the algorithm, store the values in a data block with the same bit width, and delete leading zeros to achieve the purpose of compression.

[0003] However, during the process of bit padding for a sequence, once some extreme outliers appear, it will cause the bit width of the entire sequence to increase excessively, resulting in an increase in the time and space costs of data compression. Summary of the Invention

[0004] The present invention provides a data compression method, device, electronic device and storage medium based on anomaly separation, which is used to solve the defect in the prior art that outliers increase the time and space costs of data compression. According to the normal values and outliers in the data sequence to be compressed, the data sequence to be compressed is divided into at least two data sequences, and a method of separately compressing the at least two data sequences is adopted. The effect of separating and compressing outliers for storage is achieved, the time and space costs of data compression are reduced, and the data compression efficiency is improved.

[0005] The present invention provides a data compression method based on anomaly separation, including the following steps: According to the normal values and outliers in the data sequence to be compressed, the data sequence to be compressed is divided into at least two data sequences. One of the at least two data sequences contains the normal values, and the other data sequences contain the outliers. The outliers are one or more values whose value sizes significantly deviate from the value sizes of most of the values in the data sequence to be compressed, and the number of normal values is greater than the number of outliers; Separate data compression is performed on the at least two data sequences.

[0006] According to the data compression method based on anomaly separation provided by the present invention, the step of dividing the data sequence to be compressed into at least two data sequences according to the normal values and outliers in the data sequence to be compressed includes: Determine a first boundary value and a second boundary value for partitioning the data sequence to be compressed according to the target features of each value in the data sequence to be compressed, where the target features include value size and / or bit width size, the first boundary value is less than the second boundary value, and the bit width size is used to indicate the number of position storage bits required for storing the binary value of the numerical value; Partition the data sequence to be compressed into an upper outlier sequence, a normal value sequence, and a lower outlier sequence according to the first boundary value and the second boundary value. The upper outlier sequence includes the numerical values in the data sequence to be compressed whose values are less than the first boundary value, the lower outlier sequence includes the numerical values in the data sequence to be compressed whose values are greater than the second boundary value, and the normal value sequence includes the numerical values in the data sequence to be compressed whose values are greater than or equal to the first boundary value and less than or equal to the second boundary value.

[0007] According to a data compression method based on anomaly separation provided by the present invention, where the target feature includes value size, the determining a first boundary value and a second boundary value for partitioning the data sequence to be compressed according to the target features of each value in the data sequence to be compressed includes: Sort the numerical values in the data sequence to be compressed according to value size to obtain a sorted sequence; Traverse and select different numerical values as the first boundary value and the second boundary value in the sorted sequence to obtain multiple groups of boundary value combinations, and calculate the cost corresponding to each group of boundary value combinations, where the cost is used to characterize the compression storage cost; Determine the numerical values in the boundary value combination with the smallest corresponding cost as the first boundary value and the second boundary value.

[0008] According to a data compression method based on anomaly separation provided by the present invention, the traversing and selecting different numerical values as the first boundary value and the second boundary value in the sorted sequence to obtain multiple groups of boundary value combinations and calculating the cost corresponding to each group of boundary value combinations includes: Sequentially determine the numerical values in the sorted sequence as the first boundary value in ascending order of value; After determining the first boundary value, then sequentially determine the numerical values in the sorted sequence as the second boundary value in ascending order of value and with a value greater than the first boundary value, and calculate the corresponding cost to obtain the cost corresponding to each group of boundary value combinations.

[0009] According to a data compression method based on anomaly separation provided by the present invention, where the target feature includes bit width size, the determining a first boundary value and a second boundary value for partitioning the data sequence to be compressed according to the target features of each value in the data sequence to be compressed includes: Determine the bit width of each numerical value in the data sequence to be compressed; Determine the first boundary value and the second boundary value according to the bit-width distribution corresponding to the bit-widths of the respective values.

[0010] According to a data compression method based on anomaly separation provided by the present invention, the determining the first boundary value and the second boundary value according to the bit-width distribution corresponding to the bit-widths of the respective values includes: Traverse each bit-width within the first bit-width range, and then traverse and select all the values in the data sequence to be compressed as normal values, and determine the target bit-width and the starting value corresponding to the normal values. The target bit-width is used to indicate the optimal bit-width for representing the normal values in the data sequence to be compressed, and the starting value is used to indicate the optimal starting value for representing the normal values in the data sequence to be compressed; Based on the target bit-width and the starting value, traverse and select all the values in the data sequence to be compressed as the starting values of the normal values, and then traverse and select the anomaly values and calculate the corresponding cost, where the cost is used to characterize the compression storage cost; Determine the value with the minimum corresponding cost as the first boundary value and the second boundary value.

[0011] According to a data compression method based on anomaly separation provided by the present invention, the target features include value size and bit-width size. The determining the first boundary value and the second boundary value for dividing the data sequence to be compressed according to the target features of the respective values in the data sequence to be compressed includes: Obtain the median of the data sequence to be compressed; Traverse and select normal values centered on the median of the data sequence to be compressed; Determine the bit-width range of the normal values, and determine the first boundary value and the second boundary value according to the bit-width range of the normal values.

[0012] According to a data compression method based on anomaly separation provided by the present invention, the obtaining the median of the data sequence to be compressed includes: Perform differential processing on each value in the data sequence to be compressed to obtain a processed data sequence; Observe the normal distribution of the respective values in the processed data sequence; Determine the median according to the normal distribution of the respective values.

[0013] According to a data compression method based on anomaly separation provided by the present invention, the at least two data sequences include a target data sequence, and the target data sequence includes the normal values. The respectively performing data compression on the at least two data sequences includes: Determine the minimum bit-width required to store the normal values according to the number of the normal values; The normal value is stored with the minimum bit width, and the first value stored with the minimum bit width is the starting value in the normal value, and the last value stored with the minimum bit width is the ending value in the normal value.

[0014] The present invention also provides a data compression device based on anomaly separation, including the following modules: A data partitioning module, configured to partition the data sequence to be compressed into at least two data sequences according to the normal values and anomaly values in the data sequence to be compressed, one of the at least two data sequences includes the normal values, and the other data sequences include the anomaly values, where the anomaly values are one or more values whose value sizes significantly deviate from the value sizes of most of the values in the data sequence to be compressed, and the number of the normal values is greater than the number of the anomaly values; A data compression module, configured to perform data compression on the at least two data sequences respectively.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, the data compression method based on anomaly separation as described in any one of the above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data compression method based on anomaly separation as described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the data compression method based on anomaly separation as described in any one of the above is implemented.

[0018] The data compression method, device, electronic device, and storage medium based on anomaly separation provided by the present invention partition the data sequence to be compressed into at least two data sequences according to the normal values and anomaly values in the data sequence to be compressed, one of the at least two data sequences includes the normal values, and the other data sequences include the anomaly values; and perform data compression on the at least two data sequences respectively. By separating the normal values and anomaly values in the data sequence to be compressed and performing data compression respectively, this method can reduce the time and space costs of data compression and improve the data compression efficiency. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart of the data compression method based on abnormal separation provided by the present invention.

[0021] Figure 2 It is a schematic flowchart of the method for dividing the data sequence to be compressed into at least two data sequences provided by the present invention.

[0022] Figure 3 It is an example of separating and dividing sequence outliers according to the exact value size provided by the present invention.

[0023] Figure 4 It is a schematic diagram of sequence sorting provided by the present invention.

[0024] Figure 5 It is a schematic diagram of the scenario for dividing outliers provided by the present invention.

[0025] Figure 6 It is a schematic diagram of cumulative counting provided by the present invention.

[0026] Figure 7 It is a schematic flowchart of the method for data compression of at least two data sequences provided by the present invention.

[0027] Figure 8 It is a schematic structural diagram of the data compression device based on abnormal separation provided by the present invention.

[0028] Figure 9 It is a schematic physical structure diagram of the electronic device provided by the present invention. Detailed implementation manners

[0029] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0030] The Internet of Things and time series analysis are two rapidly developing fields that play important roles in many applications. With the popularization of Internet of Things devices, the types and quantities of time series data have witnessed an explosive growth, making the efficient compression and storage of time series data increasingly important. In the work of compressing data values by traditional compression algorithms, bit stuffing is a classic task. Many classic algorithms use bit stuffing as a component in the algorithm, storing the values in a data block with the same bit width and removing leading zeros to achieve the purpose of compression.

[0031] However, during the process of sequence bit stuffing, once some extreme outliers appear, it will cause the bit width of the entire sequence to increase excessively, resulting in an increase in the time and space costs of data compression.

[0032] In view of this, an embodiment of the present invention provides a data compression method based on outlier separation. According to the normal values and outliers in the data sequence to be compressed, the data sequence to be compressed is divided into at least two data sequences. One of the at least two data sequences contains the normal values, and the other data sequences contain the outliers; the at least two data sequences are respectively compressed. This method separates the normal values and outliers in the data sequence to be compressed and compresses them respectively, which can reduce the time and space costs of data compression and improve the data compression efficiency.

[0033] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.

[0034] Figure 1 It is a schematic flowchart of the data compression method based on outlier separation provided by the present invention. The data compression method based on outlier separation can be applied to an electronic device, and this electronic device can be various types of devices with information processing capabilities during implementation. For example, the electronic device can include a personal computer, a laptop, a handheld computer, or a server, etc.; the electronic device can also be a mobile terminal. For example, the mobile terminal can include a mobile phone, an in-vehicle computer, a tablet computer, or a projector, etc. As Figure 1 shown, this method can include the following steps 101 to step 102: Step 101: According to the normal values and outliers in the data sequence to be compressed, the data sequence to be compressed is divided into at least two data sequences. One of the at least two data sequences contains the normal values, and the other data sequences contain the outliers. The outliers are used to indicate one or more values whose value sizes deviate significantly from the value sizes of most of the values in the data sequence to be compressed, and the number of normal values is greater than the number of outliers.

[0035] It should be noted that determining the normal values and abnormal values in the data sequence to be compressed and dividing the data sequence to be compressed into at least two data sequences, where the normal values form one data sequence and the abnormal values can form at least one data sequence. This step lays the foundation for subsequent separate compression. Among them, the number of determined normal values is greater than the number of abnormal values, which can reduce the storage cost. An abnormal value refers to one or more values whose magnitude significantly deviates from the magnitude of most values in the data sequence to be compressed.

[0036] Among them, there are many ways to divide the data sequence to be compressed into at least two data sequences according to the normal values and abnormal values in the data sequence to be compressed. For example, the normal values and abnormal values in the data sequence to be compressed can be determined according to a preset normal value range or a statistically determined normal value range, and then the data sequence to be compressed can be divided. The present invention does not limit the way of dividing the data sequence to be compressed into at least two data sequences according to the normal values and abnormal values in the data sequence to be compressed.

[0037] Step 102: Perform data compression on the at least two data sequences respectively.

[0038] It should be noted that there are many ways to perform data compression on the at least two data sequences respectively. For example, it can be to determine the bit width required for each value in each data sequence and then store it with the corresponding bit width, or it can be to store it according to a preset bit width, etc. The present invention does not limit the way of performing data compression on the at least two data sequences respectively.

[0039] It can be understood that the method provided by the present invention for separately compressing and storing normal values and abnormal values after separation, that is, improving the bit stuffing algorithm based on anomaly separation for data processing, realizes a more efficient data compression method.

[0040] In some embodiments, dividing the data sequence to be compressed into at least two data sequences can divide the data sequence to be compressed into three data sequences, that is, determining two boundary values and dividing the data sequence to be compressed into an upper abnormal value sequence, a normal value sequence, and a lower abnormal value sequence.

[0041] Figure 2 is a schematic flowchart of the method for dividing the data sequence to be compressed into at least two data sequences provided by the present invention. As Figure 2 shown, dividing the data sequence to be compressed into at least two data sequences according to the normal values and abnormal values in the data sequence to be compressed may include: Step 201: Determine a first boundary value and a second boundary value for partitioning the data sequence to be compressed according to the target features of each value in the data sequence to be compressed. The target features include value size and / or bit width size. The first boundary value is less than the second boundary value. The bit width size is used to indicate the number of position storage bits required for storing the binary value of the numerical value.

[0042] It should be noted that determining the first boundary value and the second boundary value for partitioning the data sequence to be compressed according to the target features of each value in the data sequence to be compressed specifies that the implementation scheme of anomaly separation is not a simple heuristic, but the first boundary value and the second boundary value determined according to the corresponding value and / or position storage cost.

[0043] Step 202: Partition the data sequence to be compressed into an upper outlier sequence, a normal value sequence, and a lower outlier sequence according to the first boundary value and the second boundary value. The upper outlier sequence includes the numerical values in the data sequence to be compressed whose values are less than the first boundary value. The lower outlier sequence includes the numerical values in the data sequence to be compressed whose values are greater than the second boundary value. The normal value sequence includes the numerical values in the data sequence to be compressed whose values are greater than or equal to the first boundary value and less than or equal to the second boundary value.

[0044] It should be noted that after determining the first boundary value and the second boundary value, the data sequence to be compressed can be partitioned into an upper outlier sequence, a normal value sequence, and a lower outlier sequence, so as to perform separate compression storage.

[0045] It can be understood that since the data sequence to be compressed is partitioned into an upper outlier sequence, a normal value sequence, and a lower outlier sequence, in the subsequent storage process of the outlier sequence and the normal value sequence, the outlier sequence needs to store some additional position information about them, and the cost is higher than that of the normal value sequence. The normal value sequence will be stored in a more convenient, direct, and compact way. Therefore, determining the first boundary value and the second boundary value for partitioning the data sequence to be compressed according to the value size and / or bit width size of each numerical value in the data sequence to be compressed can improve the accuracy of sequence partitioning and reduce the overall compression storage cost of the data sequence to be compressed.

[0046] In some embodiments, there are three methods for determining the first boundary value and the second boundary value for partitioning the data sequence to be compressed according to the target features of each value in the data sequence to be compressed, namely partitioning according to value size, partitioning according to bit width size, and partitioning according to value size and bit width size. The partitioning method mainly involves a two-dimensional search, that is, finding two boundary numerical values, that is, selecting different ways to traverse and search, which will be introduced in detail below.

[0047] Method 1: Separate according to the exact value size to partition the sequence outliers.

[0048] In an embodiment of the present invention, the target feature includes a value size. Determining a first boundary value and a second boundary value for partitioning the data sequence to be compressed according to the target features of each numerical value in the data sequence to be compressed may include: sorting the numerical values in the data sequence to be compressed according to the value size to obtain a sorted sequence; traversing and selecting different numerical values as the first boundary value and the second boundary value in the sorted sequence to obtain multiple groups of boundary value combinations, and calculating the cost corresponding to each group of boundary value combinations, where the cost is used to characterize the compression storage cost; determining the numerical values in the boundary value combination with the smallest corresponding cost as the first boundary value and the second boundary value.

[0049] It should be noted that the method of separating outliers according to exact values to partition the outlier sequence is the most accurate partitioning method. It considers all possible values in the data sequence to be compressed as the first boundary value and the second boundary value, and uses a relatively long search time. This method ensures the correctness of the separation and can obtain the minimum storage cost.

[0050] Further, traversing and selecting different numerical values as the first boundary value and the second boundary value in the sorted sequence to obtain multiple groups of boundary value combinations and calculating the cost corresponding to each group of boundary value combinations may include: sequentially determining the numerical values in the sorted sequence as the first boundary value in ascending order of value; after determining the first boundary value, then sequentially determining the numerical values in the sorted sequence as the second boundary value in ascending order of value and greater than the first boundary value, and calculating the corresponding cost to obtain the cost corresponding to each group of boundary value combinations.

[0051] Exemplarily, taking a time series (0, 124, 126, 125, 127, 128, 131, 130, 128, 255) containing 10 numbers as an example, the smallest is 0 and the largest is 255. First, sort to obtain the data sequence (0, 124, 125, 126, 127, 128, 128, 130, 131, 255), and then traverse all the values to find the optimal boundary method. For example, the small boundary is taken as 0, and the large boundary traverses 124, 125, 126... According to the preliminary expression introduced below, the cost of each boundary selection can be calculated. After traversing all possibilities, compare and select the boundary value corresponding to the minimum cost. This traversal method is a two-layer loop, so the complexity is n squared.

[0052] Exemplarily, the preliminary expression can be: Define 1, the storage cost of bit packing: (1), where is The maximum and minimum values in the sequence represent the length of the sequence.

[0053] In the separation algorithm, the key concern is to find two boundary values , with the help of which the original sequence can be divided into three types of values: Definition 2, normal value: (2), Definition 3, lower outlier: (3), Definition 4, upper outlier: (4).

[0054] At this time, the storage cost of bit-packing with outlier separation is divided into three parts, and a fourth part for storing the positions of outliers, each part being similar to that in Definition 1.

[0055] Definition 5, storage cost of bit-packing with outlier separation: Among them, represents the length of the lower outlier, represents the length of the upper outlier.

[0056] Definition 6, cumulative count, records the number of numbers in the sequence less than and less than or equal to : (6) Definition 7, based on Definition 5 and Definition 6, the storage cost of bit-packing with outlier separation can be re-expressed as: It can be understood that the method of separating outliers according to the exact value size ensures the correctness of the separation and can obtain the minimum storage cost.

[0057] Exemplarily, the pseudocode of Method 1 is as follows: , that is, the most naive way to traverse and find the optimal two boundary values .

[0058] Among them, the input is the sequence , and the output is the optimal solution of the first boundary value and the second boundary value. Step 1 represents sorting the sequence, Step 2 represents traversing the X sequence orderly, Step 3 represents obtaining the count value based on Definition 6, Step 4 represents calculating a basic storage cost, and Step 5 represents traversing the X sequence orderly again as The selected value, step 6 means traversing the second half of the X sequence in order as The selected value, step 7 means calculating the storage cost of this selection based on Definition 7, step 8 means comparing this cost with the optimal boundary scheme in the record, and step 9 means updating the lower boundary value in the optimal storage scheme , step 10 means updating the lower boundary value in the optimal storage scheme , step 11 means updating the storage cost in the optimal storage scheme, and step 12 means returning the optimal boundary value selection .

[0059] Figure 3 is an example provided by the present invention for separating and dividing sequence outliers according to the exact value size. As Figure 3 shown, represents that the size range of the upper outlier is less than , represents that the size range of the normal value is less than , represents that the size range of the lower outlier is less than .

[0060] Figure 4 is a schematic diagram of sequence sorting provided by the present invention. As Figure 4 shown, first, the present invention sorts the sequence in ascending order and obtains the cumulative count as shown in Figure 4 . In this sequence, Algorithm 1 enumerates from the minimum value 465 to the maximum value 935 , from the next value to the maximum value 935 and enumerates . Finally, the algorithm finds the optimal solution (632, 696) with the minimum cost.

[0061] Method 2, separating according to the exact bit width size to divide sequence outliers.

[0062] In the embodiments of the present invention, the target feature includes the bit width size. Determining the first boundary value and the second boundary value for dividing the data sequence to be compressed according to the target features of each value in the data sequence to be compressed may include: determining the bit width of each value in the data sequence to be compressed; determining the first boundary value and the second boundary value according to the bit width distribution corresponding to the bit width of each value.

[0063] It should be noted that based on the method of determining the segmentation boundary according to the exact value size, the present invention proposes an effective separation method based on the bit width. In this scheme, based on the bit width as the separator, the boundary value can be determined with less search time.

[0064] It can be understood that by using the above method of separating according to the exact bit width size to divide sequence outliers, the optimal solution of bit width separation can be obtained as the basis for data sequence separation, improving the division efficiency.

[0065] Further, determining the first boundary value and the second boundary value according to the bit width distribution corresponding to the bit widths of the respective numerical values may include: traversing each bit width within the first bit width range, then traversing and selecting all the numerical values in the data sequence to be compressed as normal values, determining the target bit width and the starting numerical value corresponding to the normal values, where the target bit width is used to indicate the optimal bit width for representing the normal values in the data sequence to be compressed, and the starting numerical value is used to indicate the optimal starting numerical value for representing the normal values in the data sequence to be compressed; based on the target bit width and the starting numerical value, traversing and selecting all the numerical values in the data sequence to be compressed as the starting numerical values of the normal values, then traversing and selecting outliers and calculating the corresponding cost, where the cost is used to characterize the compression storage cost; and determining the numerical value with the minimum corresponding cost as the first boundary value and the second boundary value.

[0066] It should be noted that the method of separating according to the exact bit width size to divide sequence outliers includes traversing the bit width range to find the optimal segmentation.

[0067] Exemplarily, taking a time series (0, 124, 126, 125, 127, 128, 131, 130, 128, 255) containing 10 numbers as an example. Two rounds of traversal are required. In the first round, first traverse all possible bit widths of the normal values (finally select 3 bits), and then traverse and select all values as possible starting points of the normal values (124). In the second round, first traverse and select the starting points of the normal values, and then traverse the upper outliers, that is, the possible bit widths of the overly large values. The optimal separation scheme is selected during the two rounds of traversal. Such complexity is n logn.

[0068] It can be understood that the precise bit width separation method provided by the present invention only considers the specific bit widths of the central value and the upper outliers within a short search time, improving the search efficiency. By explaining all possible search situations, the bit width separation still returns the optimal solution as the value separation. Its correctness is ensured by transforming the solution based on the values from the sequence value set.

[0069] Exemplarily, the pseudo code for separating according to the exact bit width size to divide sequence outliers is as follows: Among them, the input is the sequence , and the output is the optimal solution of the first boundary value and the second boundary value. Step 1 represents Sort the sequence. Steps 2 to 3 represent traversing the X sequence in order and obtaining the count value based on Definition 6. Step 4 represents calculating a basic storage cost. Steps 5 to 6 represent traversing the X sequence again in order as the selected value. Steps 7 to 8 represent traversing to determine the lower boundary value selection. Step 9 represents calculating the storage cost of this selection based on Definition 7. Steps 10 to 13 represent comparing this cost with the optimal boundary scheme in the record and updating accordingly. Steps 14 to 15 represent traversing to determine the lower boundary value selection. Step 16 represents calculating the storage cost of this selection based on Definition 7. Steps 17 to 20 similarly represent comparing this cost with the optimal boundary scheme in the record and updating accordingly. Step 21 represents returning the optimal boundary value selection .

[0070] First, through relatively complex mathematical discussions, it is found that only the separation cases shown in the following table need to be summarized: Table I All possible cases separated according to bit width

[0071] Among them, similar to Algorithm 1, the present invention calculates the cumulative count of values in lines 1 - 3. Then, this algorithm enumerates each and the cost of each corresponding β, where β ≤ γ. According to the first case in Table II, the solution to be considered is ( , min + ). For the second case in Table II, the algorithm enumerates each under the solution in lines 14 - 20 and the cost of each γ.

[0072] Still consider the sequence in the previous one. First, we sort the sequence in ascending order and obtain the cumulative count, similar to Algorithm 1. In this sequence, Algorithm 2 enumerates from the minimum value 465 to the maximum value 935. For each = and β ≤ γ, that is, , we consider the cost of = + .

[0073] Figure 5 is a schematic diagram of the scenario for partitioning sequence outliers provided by the present invention. As Figure 5 shown, for each = and β > γ, that is , we consider , as shown in the figure below. Finally, the algorithm found = 632 and the optimal solution with β = 6.

[0074] Method 3: Separate and divide sequence outliers according to the exact value size and bit width size.

[0075] In an embodiment of the present invention, the target feature includes a value size and a bit width size. The determining of the first boundary value and the second boundary value for dividing the to-be-compressed data sequence according to the target features of each numerical value in the to-be-compressed data sequence may include: obtaining the median of the to-be-compressed data sequence; traversing and selecting normal values centered on the median of the to-be-compressed data sequence; determining the bit width range of the normal values, and determining the first boundary value and the second boundary value according to the bit width range of the normal values.

[0076] It should be noted that the present invention proposes a strategy for running approximate median separation in a very short time. By using the above method, outlier separation can be completed in a relatively short time.

[0077] Further, the obtaining of the median of the to-be-compressed data sequence may include: performing a difference process on each numerical value in the to-be-compressed data sequence to obtain a processed data sequence; observing the normal distribution of each numerical value in the processed data sequence; and determining the median according to the normal distribution of each numerical value.

[0078] It should be noted that the method of using an approximate median separation algorithm to divide sequence outliers is similar to the method of separating based on bit width size. Assuming that the selection of normal values is only centered on the median of the sequence, the bit width range of the normal values is searched, and those outside are outliers. At this time, only one-dimensional search is required, with a complexity of O(n).

[0079] Exemplarily, taking a time series containing 10 numbers (0, 124, 126, 125, 127, 128, 131, 130, 128, 255) as an example. In the example, it can be understood that the search is around 128.

[0080] It can be understood that the present invention proposes an approximate median separation method. By observing the normal distribution of values, especially after difference processing, the median is considered in the separation. Combining the above bit widths, approximate separation can be determined in a very short time.

[0081] Exemplarily, the pseudo-code of the method for separating and dividing sequence outliers according to the exact value size and bit width size is as follows: Among them, the input is a sequence , and the output is the optimal solution of the first boundary value and the second boundary value. Step 1 means finding the median value of the sequence, step 2 means traversing and selecting all values in the sequence, step 3 means that 10 is a branch structure, and according to the size relationship between the selected value and the median of the sequence, the count of different buckets is increased. Step 11 means calculating a basic storage cost, and step 12 means traversing and selecting the bit width of the normal value , steps 13 to 14 mean updating the count of the upper and lower outlier numbers in the loop, and steps 15 to 16 mean based on the bit width of the normal value determined upper and lower boundary values. Step 17 determines the storage cost at this time based on Definition 5. Steps 18 to 21 mean comparing this cost with the optimal boundary scheme in the record and updating accordingly. Step 22 means returning the optimal boundary value selection .

[0082] Figure 6 is a schematic diagram of the cumulative count provided by the present invention. As Figure 6 shown, still consider the sequence in the previous one. First, still consider the previous sequence. First, Algorithm 3 finds the median = 674. Given the maximum value β = 9, it divides the values into 19 buckets. Then, the algorithm searches for the bit width β with the minimum cost by enumerating β from 9 to 1. It returns the solution (610, 738) with β = 6. As Figure 6 shown, only need to widen the ranges of several integer powers of 2 from the median to both sides and select them

[0083] In summary, Method 1 is implemented in a relatively direct and rough way. It traverses and searches for the upper and lower bounds of outliers to perform the division, and can obtain the most accurate division result. Method 2 is optimized. It traverses the bit widths used for upper and lower outliers to complete the division, and the speed is much improved. At the same time, the compression ratio remains unchanged. Method 3 is an approximate algorithm. The compression rate is not as good as the second one but the difference is not large. At the same time, the compression speed is much faster, and the shortest compression and decompression time can be obtained

[0084] In some embodiments, after determining the normal value and obtaining the normal value sequence, the minimum bit width can be determined for compression according to the normal value

[0085] Figure 7 is a schematic flowchart of a method for data compression of at least two data sequences provided by the present invention. As Figure 7 shown, the at least two data sequences include a target data sequence, the target data sequence includes the normal value, and respectively performing data compression on the at least two data sequences may include: Step 301: Determine the minimum bit width required to store the normal value according to the range of the normal value.

[0086] Step 302: Store the normal value with the minimum bit width. The first value stored with the minimum bit width is the minimum value in the normal value, and the last value stored with the minimum bit width is the maximum value in the normal value.

[0087] It should be noted that after determining the range of each sequence, it can be stored one by one using the minimum bit width, which is also part of the design in the system.

[0088] Exemplarily, taking a time series including 10 numbers (0, 124, 126, 125, 127, 128, 131, 130, 128, 255) as an example, the minimum is 0 and the maximum is 255. Originally, each number would require 8 bits (2^8 = 256) to store. However, in fact, most (8) normal values are in the range of 124 - 131, with only 8 possible values, and only 3 bits (2^3) are needed to store. That is, 000 represents 124, 001 represents 125, 010 represents 126, 011 represents 127... In this way, three binary bits are used to store each normal value. This corresponding relationship only needs to store the boundary values of the upper and lower bounds. At the same time, the range of outliers can also be determined in a smaller range, thus also using a smaller bit width for storage.

[0089] It can be understood that after determining the outliers, the present invention proposes an optimized bit - stuffing data compression method. Specifically, the present invention proposes bit - stuffing based on outlier separation, which improves bit - stuffing by separating the upper outlier sequence and the lower outlier sequence. Using this method can further reduce the compression storage cost.

[0090] The method provided by the present invention for separately storing the upper and lower outlier sequences, although the remaining central values have a narrow expansion, that is, the compressed bit width, the separated outliers require some additional cost to represent their positions. Therefore, a better threshold is used to separate the upper and lower outliers, thereby generating a smaller storage cost. Instead of enumerating all possible values as the upper and lower outlier separators, which would take too long.

[0091] It can be understood that the data compression algorithm based on outlier separation of the present invention is compatible with any existing compression method using bit - packing, and can replace the bit - stuffing algorithm block in multiple algorithms after deployment. A large number of experiments on many real - world datasets show that among various compression methods, the compression ratio of the replaced data compression algorithm is significantly increased from about 2.75 to 3.25.

[0092] Next, the exemplary application of the embodiments of the present invention in an actual application scenario will be described.

[0093] Take Figure 3 as an example. As Figure 3 shown, Xmax represents the upper outlier, Xu represents the normal value, Xmin represents the lower outlier, represents the first boundary value, represents the second boundary value. This figure is obtained by the outlier separation method, which reflects the method of determining outlier separation according to two boundary values, and then guides the subsequent bit packing and compression.

[0094] Based on Figure 3 , the data sequence to be compressed is divided into three parts: the normal value sequence, the upper / lower outlier sequence. The range of a large number of normal values is restricted to a relatively small range, so that the bit width required to store them can be greatly reduced.

[0095] In addition, in order to verify the data compression method based on outlier separation provided by the present invention, the present invention implements the deployment of the data compression method based on outlier separation, and conducts extensive experiments on many real-world data sets. The experiments show that the compression ratio is significantly increased from about 2.75 of the existing method to an average of 3.25 of the exact bit width separation method. Therefore, it is strongly recommended to replace the simple bit filling with the data compression method based on outlier separation proposed by the present invention in practice.

[0096] Based on the foregoing embodiments, an embodiment of the present invention provides a data compression device based on outlier separation. Each module included in the device, and each unit included in each module, can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the process of implementation, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0097] The data compression device based on outlier separation provided by the present invention will be described below. The data compression device based on outlier separation described below can be mutually referred to the data compression method based on outlier separation described above.

[0098] Figure 8 is a schematic structural diagram of the data compression device based on outlier separation provided by the present invention. As Figure 8 shown, the device 500 includes a data division module 501 and a data compression module 502, wherein: A data division module 501 is configured to divide the to-be-compressed data sequence into at least two data sequences according to normal values and abnormal values in the to-be-compressed data sequence. One of the at least two data sequences contains the normal values, and the other data sequences contain the abnormal values. The abnormal values are one or more values whose value sizes significantly deviate from the value sizes of most of the values in the to-be-compressed data sequence. The number of normal values is greater than the number of abnormal values. A data compression module 502 is configured to perform data compression on the at least two data sequences respectively.

[0099] In some embodiments, the data division module 501 includes a boundary determination unit and a data division unit, where the boundary determination unit is configured to determine a first boundary value and a second boundary value for dividing the to-be-compressed data sequence according to the target features of each value in the to-be-compressed data sequence. The target features include value size and / or bit-width size. The first boundary value is less than the second boundary value. The bit-width size is used to indicate the number of position storage bits required to store the binary value of the numerical value. the data division unit is configured to divide the to-be-compressed data sequence into an upper abnormal value sequence, a normal value sequence, and a lower abnormal value sequence according to the first boundary value and the second boundary value. The upper abnormal value sequence includes the values in the to-be-compressed data sequence that are less than the first boundary value. The lower abnormal value sequence includes the values in the to-be-compressed data sequence that are greater than the second boundary value. The normal value sequence includes the values in the to-be-compressed data sequence that are greater than or equal to the first boundary value and less than or equal to the second boundary value.

[0100] In some embodiments, the target feature includes value size, and the boundary determination unit includes a data sorting component, a cost calculation component, and a first determination component, where the data sorting component is configured to sort the values in the to-be-compressed data sequence according to value size to obtain a sorted sequence; the cost calculation component is configured to traverse and select different values as the first boundary value and the second boundary value in the sorted sequence to obtain multiple groups of boundary value combinations, and calculate the cost corresponding to each group of boundary value combinations. The cost is used to characterize the compression storage cost; the first determination component is configured to determine the values in the boundary value combination with the minimum corresponding cost as the first boundary value and the second boundary value.

[0101] In some embodiments, the cost calculation component is specifically configured to: sequentially determine the numerical values in the sorted sequence as the first boundary values in ascending order of value; after determining the first boundary values, sequentially determine the numerical values in the sorted sequence as the second boundary values in ascending order of value and greater than the first boundary values, and calculate the corresponding costs to obtain the costs corresponding to each group of boundary value combinations.

[0102] In some embodiments, the target feature includes the bit width size, and the boundary determination unit includes a bit width determination component and a second determination component, where the bit width determination component is configured to determine the bit width of each numerical value in the data sequence to be compressed; the second determination component is configured to determine the first boundary value and the second boundary value according to the bit width distribution corresponding to the bit width of each numerical value.

[0103] In some embodiments, the second determination component is specifically configured to: traverse each bit width within the first bit width range, then traverse and select all the numerical values in the data sequence to be compressed as normal values, and determine the target bit width and the starting numerical value corresponding to the normal values, where the target bit width is used to indicate the optimal bit width for representing the normal values in the data sequence to be compressed, and the starting numerical value is used to indicate the optimal starting numerical value for representing the normal values in the data sequence to be compressed; Based on the target bit width and the starting numerical value, traverse and select all the numerical values in the data sequence to be compressed as the starting numerical values of the normal values, then traverse and select the abnormal values and calculate the corresponding costs, where the costs are used to characterize the compression storage cost; Determine the numerical values with the minimum corresponding costs as the first boundary value and the second boundary value.

[0104] In some embodiments, the target feature includes the value size and the bit width size, and the boundary determination unit includes a median determination component, a normal value determination component, and a third determination component, where the median determination component is configured to obtain the median of the data sequence to be compressed; the normal value determination component is configured to traverse and select normal values centered around the median of the data sequence to be compressed; the third determination component is configured to determine the bit width range of the normal values, and determine the first boundary value and the second boundary value according to the bit width range of the normal values.

[0105] In some embodiments, the median determination component is specifically configured to: perform a difference processing on each numerical value in the data sequence to be compressed to obtain a processed data sequence; observe the normal distribution of each numerical value in the processed data sequence; and determine the median according to the normal distribution of each numerical value.

[0106] In some embodiments, the at least two data sequences include a target data sequence, and the target data sequence includes the normal values. Specifically, the data compression module is configured to: determine the minimum bit width required to store the normal values according to the number of the normal values; store the normal values with the minimum bit width, where the first value stored with the minimum bit width is the starting value among the normal values, and the last value stored with the minimum bit width is the ending value among the normal values.

[0107] In the embodiments of the present invention, a method for data compression is provided, which can divide a data sequence to be compressed into at least two data sequences according to normal values and abnormal values in the data sequence to be compressed, and perform data compression on the at least two data sequences respectively. The effect of separating and compressing abnormal values for storage is achieved, the time and space costs of data compression are reduced, and the data compression efficiency is improved.

[0108] Figure 9 is a schematic physical structure diagram of an electronic device provided by the present invention. As Figure 9 shown, the electronic device 600 may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a data compression method based on anomaly separation, and the method includes: dividing the data sequence to be compressed into at least two data sequences according to normal values and abnormal values in the data sequence to be compressed, where one of the at least two data sequences contains the normal values, and the other data sequences contain the abnormal values. The abnormal values are one or more values whose value sizes significantly deviate from the value sizes of most of the values in the data sequence to be compressed, and the number of the normal values is greater than the number of the abnormal values; performing data compression on the at least two data sequences respectively.

[0109] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0110] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data compression method based on anomaly separation provided by the above-mentioned various methods. The method includes: dividing the data sequence to be compressed into at least two data sequences according to the normal values and anomaly values in the data sequence to be compressed, one of the at least two data sequences contains the normal values, and the other data sequences contain the anomaly values. The anomaly values are used to indicate one or more values whose value sizes significantly deviate from the value sizes of most of the values in the data sequence to be compressed, and the number of normal values is greater than the number of anomaly values; respectively performing data compression on the at least two data sequences.

[0111] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are all or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0112] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a data compression method based on anomaly separation provided by the above-mentioned various methods. The method includes: dividing the data sequence to be compressed into at least two data sequences according to the normal values and anomaly values in the data sequence to be compressed, where one of the at least two data sequences contains the normal values and the other data sequences contain the anomaly values. The anomaly values are one or more values used to indicate that the value size significantly deviates from the value size of most of the values in the data sequence to be compressed, and the number of normal values is greater than the number of anomaly values; and separately performing data compression on the at least two data sequences.

[0113] The above computer-readable storage medium may adopt any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0114] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0115] The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including - but not limited to - wireless, wire, optical cable, radio frequency (RF), etc., or any suitable combination of the above.

[0116] Computer program code for performing the operations of this specification can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A data compression method based on anomaly separation, characterized in that: include: According to normal values ​​and abnormal values ​​in the data sequence to be compressed, the data sequence to be compressed is divided into at least two data sequences, one of the at least two data sequences contains the normal values, and the other data sequences contain the abnormal values, the abnormal values ​​are used to indicate one or more values ​​whose values ​​significantly deviate from the values ​​of most values ​​in the data sequence to be compressed, and the number of the normal values ​​is greater than the number of the abnormal values; Data compression is performed on the at least two data sequences respectively.

2. The data compression method based on anomaly separation according to claim 1, characterized in that: The step of dividing the data sequence to be compressed into at least two data sequences according to normal values ​​and abnormal values ​​in the data sequence to be compressed comprises: Determine a first boundary value and a second boundary value for dividing the data sequence to be compressed according to target features of each value in the data sequence to be compressed, wherein the target features include a value size and / or a bit width size, the first boundary value is smaller than the second boundary value, and the bit width size is used to indicate the number of position storage bits required to store the binary value of the value; The data sequence to be compressed is divided into an upper outlier value sequence, a normal value sequence and a lower outlier value sequence according to the first boundary value and the second boundary value, the upper outlier value sequence includes values ​​whose median of the data sequence to be compressed is less than the first boundary value, the lower outlier value sequence includes values ​​whose median of the data sequence to be compressed is greater than the second boundary value, and the normal value sequence includes values ​​whose median of the data sequence to be compressed is greater than or equal to the first boundary value and less than or equal to the second boundary value.

3. The data compression method based on anomaly separation according to claim 2, characterized in that: The target feature includes a value size, and determining the first boundary value and the second boundary value for dividing the data sequence to be compressed according to the target feature of each value in the data sequence to be compressed includes: Sorting the values ​​in the to-be-compressed data sequence according to the magnitude of the values ​​to obtain a sorted sequence; traversing and selecting different values ​​in the sorted sequence as the first boundary value and the second boundary value to obtain multiple groups of boundary value combinations, and calculating the cost corresponding to each group of boundary value combinations, wherein the cost is used to characterize the compression storage cost; The values ​​in the boundary value combination with the minimum corresponding cost are determined as the first boundary value and the second boundary value.

4. The data compression method based on anomaly separation according to claim 3 is characterized in that: The traversing and selecting different values ​​in the sorted sequence as the first boundary value and the second boundary value to obtain multiple groups of boundary value combinations, and calculating the cost corresponding to each group of boundary value combinations, includes: In order from small to large, the values ​​in the sorted sequence are determined as the first boundary values; After determining the first boundary value, the values ​​in the sorted sequence are determined as the second boundary values ​​in order from small to large and the values ​​are greater than the first boundary value, and the corresponding costs are calculated to obtain the costs corresponding to the combinations of the boundary values.

5. The data compression method based on anomaly separation according to claim 2, characterized in that: The target feature includes a bit width, and determining a first boundary value and a second boundary value for dividing the to-be-compressed data sequence according to the target feature of each value in the to-be-compressed data sequence includes: Determining the bit width of each value in the to-be-compressed data sequence; The first boundary value and the second boundary value are determined according to the bit width distribution corresponding to the bit width of each of the numerical values.

6. The data compression method based on anomaly separation according to claim 5, characterized in that: The determining the first boundary value and the second boundary value according to the bit width distribution corresponding to the bit width of each value includes: Traversing each bit width within the first bit width range, and then traversing and selecting all values ​​in the to-be-compressed data sequence as normal values, determining a target bit width and a starting point value corresponding to the normal value, wherein the target bit width is used to indicate an optimal bit width for representing the normal value in the to-be-compressed data sequence, and the starting point value is used to indicate an optimal starting value for representing the normal value in the to-be-compressed data sequence; Based on the target bit width and the starting value, traverse and select all values ​​in the to-be-compressed data sequence as the starting value of the normal value, then traverse and select the abnormal value and calculate the corresponding cost, where the cost is used to characterize the compression storage cost; The corresponding values ​​with the minimum cost are determined as the first boundary value and the second boundary value.

7. The data compression method based on anomaly separation according to claim 2, characterized in that: The target feature includes a value size and a bit width size, and determining a first boundary value and a second boundary value for dividing the to-be-compressed data sequence according to the target feature of each value in the to-be-compressed data sequence includes: Obtaining the median of the to-be-compressed data sequence; Taking the median of the data sequence to be compressed as the center, traverse and select normal values; A bit width range of the normal value is determined, and the first boundary value and the second boundary value are determined according to the bit width range of the normal value.

8. The data compression method based on anomaly separation according to claim 7, characterized in that: The obtaining the median of the to-be-compressed data sequence includes: Performing differential processing on each value in the to-be-compressed data sequence to obtain a processed data sequence; Observe the normal distribution of each value in the processed data sequence; The median is determined according to the normal distribution of the respective values.

9. The data compression method based on anomaly separation according to claim 1, characterized in that: The at least two data sequences include a target data sequence, the target data sequence includes the normal value, and the respectively performing data compression on the at least two data sequences includes: Determine the minimum bit width required to store the normal value according to the number of the normal values; The normal value is stored by the minimum bit width, the first value stored by the minimum bit width is the starting value of the normal value, and the last value stored by the minimum bit width is the ending value of the normal value.

10. A data compression device based on anomaly separation, characterized in that: include: A data partitioning module, configured to partition the data sequence to be compressed into at least two data sequences according to normal values ​​and abnormal values ​​in the data sequence to be compressed, wherein one of the at least two data sequences includes the normal values, and the other data sequences include the abnormal values, wherein the abnormal values ​​are used to indicate one or more values ​​whose values ​​significantly deviate from the values ​​of most values ​​in the data sequence to be compressed, and the number of the normal values ​​is greater than the number of the abnormal values; The data compression module is used to perform data compression on the at least two data sequences respectively.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the data compression method based on anomaly separation as claimed in any one of claims 1 to 9 is implemented.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data compression method based on anomaly separation as claimed in any one of claims 1 to 9 is implemented.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the data compression method based on anomaly separation as claimed in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Electronic commerce time sequence data anomaly detection method and system

    CN104915846A

  • Abnormal data identification method and device, computer equipment and storage medium

    CN117493828A

  • Time sequence data compression algorithm based on anomaly detection

    CN117749194A

  • Abnormal data processing method and device, electronic equipment and storage medium

    CN118332258A

  • Methods and systems to reduce time series data and detect outliers

    US20180365298A1