Lossless application of lossy compression of huge mass time series data
By comparing lossy compression and lossless compression methods, a suitable compression scheme was determined, which solved the problem of large storage space consumption for time-series data, and achieved significant improvement in storage efficiency and reduction in cost without affecting application performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING JUYI TECH CO LTD
- Filing Date
- 2021-08-25
- Publication Date
- 2026-05-08
AI Technical Summary
The lack of research on dynamic compression of time-series data in existing technologies leads to excessive storage space consumption and makes it impossible to choose a suitable compression method to improve storage efficiency and reduce costs.
By comparing and analyzing lossy and lossless compression on massive amounts of time-series data, a suitable compression method is determined, including data processing, decompression after lossy compression and application value analysis, load warning analysis, tiered comparison, and final compression result comparison, so as to achieve lossy compression to save storage space without affecting application performance.
Without affecting application performance, lossy compression significantly improves the compression ratio, saves storage space, and reduces storage costs, making it suitable for time-series data storage in the Internet of Things.
Smart Images

Figure CN115729906B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of Internet of Things (IoT) data applications, specifically the lossless application of lossy compression of massive amounts of time-series data. Background Technology
[0002] In modern life, the Internet of Things (IoT) is developing rapidly, with the variety and quantity of various devices increasing at an alarming rate. The amount of data generated is also experiencing explosive growth, expanding from GB to TB and then to PB levels. With continuous technological advancements, the prices of numerous devices such as mobile devices, cameras, software logs, and wireless sensors are constantly decreasing, making them increasingly accessible to the general public, leading to a rapid increase in data volume. Furthermore, the rapid development of the big data era benefits from progress in the industrial and technological fields. For example, video surveillance equipment, ubiquitous in major cities, generates data every moment, with annual output reaching PB levels. These real-time video feeds are widely used in various fields, continuously driving social progress and technological development. Simultaneously, the demands and requirements for data analysis are also increasing. Therefore, the importance of data storage is evident. How to compress this massive amount of data to improve storage efficiency and reduce storage costs is a question worthy of in-depth exploration. Time-series data refers to a series of data based on time. Generally, time-series data is a data series recording the same indicator in chronological order. By connecting data points in a quadrant with time as the horizontal axis, lines can be drawn to reveal trends, patterns, and anomalies, enabling prediction and early warning. A time-series database is used to store this data, supporting basic functions such as fast writing, persistence, and multi-dimensional aggregation queries.
[0003] As the total amount of time-series data continues to increase, the memory space it occupies also becomes larger and larger. Therefore, dynamic compression of time-series data is particularly important. Choosing a suitable dynamic compression method for time-series data is crucial. However, there is a lack of research on dynamic compression of time-series data in the current technology, so it is impossible to select and determine a suitable dynamic compression method for time-series data. Summary of the Invention:
[0004] The purpose of this invention is to provide a lossless application for lossy compression of massive amounts of time-series data in order to solve the above-mentioned problems, thus solving the problems mentioned in the background art.
[0005] To address the above problems, the present invention provides a technical solution:
[0006] Lossless applications of lossy compression for massive amounts of time-series data include the following steps:
[0007] S1. First, obtain the experimental data, and then organize and store it.
[0008] S2. Then, the experimental data after lossy compression is decompressed, and the application value of the data after lossy compression and decompression is analyzed.
[0009] S3. Next, load warning analysis is performed on the lossless compression experimental data and the experimental data after lossy compression, and the warning situation of the experimental data after lossy compression is compared with the warning situation of the original experimental data.
[0010] S4. Then, the original experimental data is sorted according to the degree of contamination. At the same time, the original experimental data is lossily compressed, and the experimental data after lossy compression is sorted again. Then, the difference between the sorting of the original experimental data and the sorting of the experimental data after lossy compression is compared.
[0011] S5. Finally, lossless compression and lossy compression methods were used to compress the obtained experimental data, and the compression results were recorded. The results after compression were compared in a comparative experiment.
[0012] S6. Incorporate lossy compressed data into lossless compressed data.
[0013] Preferably, the experimental data in step S1 is current data.
[0014] Preferably, the data application value in step S2 refers to whether the data can achieve the expected purpose when actually used.
[0015] Preferably, the specific operation of the data application value analysis in step S2 is as follows: obtain the experimental data after lossy compression and decompression, and compare and analyze it with the experimental data after compression using lossless compression method, and analyze whether the experimental data after lossy compression and decompression can achieve the same effect as the experimental data after compression using lossless compression method.
[0016] As a preferred embodiment, the specific operation of performing load early warning analysis on the original experimental data in step S3 is as follows: first, analyze the load situation of the experimental data, set the early warning load, then conduct the experiment and record the number of early warnings and the early warning time period.
[0017] Preferably, the specific operation of performing load warning analysis on the experimental data after lossy compression in step S3 is as follows: compress the experimental data using a lossy compression method, analyze the load of the experimental data after lossy compression, set a warning load, and then record the number of warnings and the warning time period.
[0018] Preferably, the specific classification in step S4 is as follows: the first tier is for low-power-consuming and polluting industries, i.e., daily power consumption is less than 10,000 kWh; the second tier is for medium-power-consuming and polluting industries, i.e., daily power consumption is between 10,000 and 1 million kWh; and the third tier is for high-power-consuming and polluting industries, i.e., daily power consumption exceeds 1 million kWh.
[0019] Preferably, the experimental results recorded in step S5 are the data compression rates.
[0020] The beneficial effects of this invention are as follows: This invention analyzes and studies the application problems of data in a massive time series database, and from the perspective of application value, performs lossy and lossless compression on massive time series data without affecting the application. At the same time, it compares the lossy compression rate and the lossless compression rate to determine the advantages and disadvantages between lossy compression and lossless compression, so as to facilitate the selection of a suitable compression method and thus achieve the purpose of reducing costs and improving the compression ratio. Attached image description:
[0021] For ease of explanation, the present invention will be described in detail below with reference to specific embodiments and accompanying drawings.
[0022] Figure 1 This is a schematic diagram illustrating the daily power consumption of users in an embodiment of the present invention;
[0023] Figure 2 This is a daily electricity consumption fluctuation diagram for the entire province in an embodiment of the present invention. Detailed implementation method:
[0024] like Figure 1-2 As shown, the specific implementation adopts the following technical solution:
[0025] Example:
[0026] Lossless applications of lossy compression for massive amounts of time-series data include the following steps:
[0027] S1. First, obtain the experimental data, and then organize and store it.
[0028] S2. Then, the experimental data after lossy compression is decompressed, and the application value of the data after lossy compression and decompression is analyzed.
[0029] S3. Next, load warning analysis is performed on the lossless compression experimental data and the experimental data after lossy compression, and the warning situation of the experimental data after lossy compression is compared with the warning situation of the original experimental data.
[0030] S4. Then, the original experimental data is sorted according to the degree of contamination. At the same time, the original experimental data is lossily compressed, and the experimental data after lossy compression is sorted again. Then, the difference between the sorting of the original experimental data and the sorting of the experimental data after lossy compression is compared.
[0031] S5. Finally, lossless compression and lossy compression methods were used to compress the obtained experimental data, and the compression results were recorded. The results after compression were compared in a comparative experiment.
[0032] S6. Integrating lossy compressed data into lossless compressed data is based on the massive scale of IoT data. Lossy compression is performed to achieve lossless application. Data compression is based on real-world scenarios. First, statistical analysis of data fluctuations is conducted. Based on customer needs, points that have no impact on the application effect are removed, and the truly effective points are retained for compression. This saves storage space and does not affect the application of the data. For example, when making power early warnings, only extreme value points are needed, and other lost points have no impact on the early warning application.
[0033] The experimental data in step S1 is current data.
[0034] In step S2, the data application value refers to whether the data can achieve the expected purpose when actually used.
[0035] The specific operation of the data application value analysis in step S2 is as follows: obtain the experimental data after lossy compression and decompression, and compare and analyze it with the experimental data after compression using lossless compression method, and analyze whether the experimental data after lossy compression and decompression can achieve the same effect as the experimental data after compression using lossless compression method.
[0036] The specific operation of performing load early warning analysis on the original experimental data in step S3 is as follows: first, analyze the load situation of the experimental data, set the early warning load, then conduct the experiment and record the number of early warnings and the early warning time period.
[0037] The specific operation of performing load warning analysis on the experimental data after lossy compression in step S3 is as follows: the experimental data is compressed using a lossy compression method, and the load situation of the experimental data after lossy compression is analyzed. At the same time, a warning load is set, and then the number of warnings and the warning time period are recorded.
[0038] The specific classifications in step S4 are as follows: the first tier is for low-power-consuming and polluting industries, i.e., daily power consumption is less than 10,000 kWh; the second tier is for medium-power-consuming and polluting industries, i.e., daily power consumption is between 10,000 and 1 million kWh; and the third tier is for high-power-consuming and polluting industries, i.e., daily power consumption exceeds 1 million kWh.
[0039] The experimental results recorded in step S5 are the data compression rates.
[0040] Example:
[0041] First, the experimental data was acquired, organized, and stored. Specifically, the current data of a certain province for one day was used for analysis, with a time interval of 1 minute.
[0042] like Figure 1 The following table shows the daily power consumption of a certain device. The fluctuations are as follows: the daily power consumption changes in a clear pattern. The peak power consumption periods are 8:00-13:30, 16:00-20:30, and 22:00-00:45, while the low power consumption periods are 00:45-08:00, 13:30-16:00, and 20:30-22:00. Among these, there are three extreme values during the peak power consumption periods and two extreme values during the low power consumption periods.
[0043] Then, the experimental data after lossy compression is decompressed, and the application value of the data after lossy compression and decompression is analyzed. Specifically, after decompressing the lossy compressed data, it can still show the high-level power consumption time, low-level power consumption time, and extreme power consumption time (such as power outage). Lossy compression can achieve the same effect as lossless compression, with the same data application value, while also saving costs.
[0044] Next, load warning analysis was performed on the lossless compression experimental data and the experimental data after lossy compression, and the warning situation of the experimental data after lossy compression was compared with that of the original experimental data. Specifically, two application scenarios in the power industry were selected for analysis.
[0045] Load warning: Large-scale power grid outages are a disaster in modern society, causing enormous losses to the national economy and severely impacting society and people's lives. Therefore, load warning is crucial for the power industry. Load warnings are issued by analyzing the province's daily electricity load and the overall peak-to-valley ratio. A warning is triggered when electricity consumption reaches 80% of the total load. Analysis of province-wide data shows that peak electricity consumption tends to occur at noon and 8 pm, both exceeding 80% of the total load. Figure 2 The table shows the daily electricity consumption fluctuation chart for the entire province. Further analysis of the provincial data was then conducted, dividing the data by region to identify transformer substations 1, 2, 3, ..., 72000, for which early warnings were issued. The number of early warnings is shown in Table 1.
[0046] Table 1 Number of Regional Warnings
[0047]
[0048] The warning periods are shown in Table 2:
[0049] Table 2 Warning Time
[0050]
[0051]
[0052] Obtain the number of warnings and the warning period for each transformer station.
[0053] After lossy compression of the original data, Tables 3 and 4 were obtained through data analysis as follows:
[0054] Table 3 Regional Early Warning
[0055]
[0056] Table 4 Warning Time
[0057]
[0058] It can be seen that the two compression methods obtained the same warning results. The number of warnings is shown in Table 1 and Table 3, and the warning time period is shown in Table 2 and Table 4.
[0059] Then, the original experimental data were sorted according to the degree of contamination. At the same time, the original experimental data were lossily compressed, and the experimental data after lossy compression were sorted again. The differences between the sorting of the original experimental data and the sorting of the experimental data after lossy compression were then compared.
[0060] As the electricity consumption of major industries increases, the tight power supply situation in the country is further aggravated, especially when encountering extreme weather, or when the external power supply capacity is limited or the unit fails. In such cases, power rationing policies are generally adopted. Therefore, power rationing for polluting enterprises can minimize the impact of power shortages on society, economy and residents' lives. This paper uses the daily current consumption data of polluting enterprises in a certain province for analysis.
[0061] Polluting enterprises in a certain province can be categorized into culture, sports and entertainment, mining, construction, scientific research and technical services, retail, wholesale, and manufacturing. Among them, the manufacturing industry has the most serious pollution situation (as shown in Table 5), accounting for 94.86% of the total electricity consumption.
[0062] Table 5 Daily Electricity Consumption of Pollution-Related Industries in a Certain Province
[0063] Industry type Daily electricity consumption (kWh) Culture, sports and entertainment industry 989154.6 Mining 146687.9 Construction industry 760785 Scientific research and technical services industry 763315.6 Retail 564452.9 Wholesale 1449.523 manufacturing 59559979
[0064] Based on daily power consumption, industries subject to pollution are divided into three tiers: Tier 1 – Low power consumption (less than 10,000 kWh per day); Tier 2 – Medium power consumption (between 10,000 and 1 million kWh per day); and Tier 3 – High power consumption (more than 1 million kWh per day). See the table below:
[0065] Table 6 Classification of Polluting Enterprises
[0066] Industry type Non-destructive enterprise classification Culture, sports and entertainment industry First Tier Mining Second Tier Construction industry Second Tier Scientific research and technical services industry Second Tier Retail Second Tier Wholesale Second Tier manufacturing Third Tier
[0067] In a certain province, the culture, sports and entertainment industry is classified as a low-energy-consuming and polluting industry; the mining, construction, scientific research and technical services, retail and wholesale industries are classified as medium-energy-consuming and polluting industries; and the manufacturing industry is classified as a high-energy-consuming and polluting enterprise.
[0068] The manufacturing sector in one city accounts for 42.87% of the province's total electricity consumption among polluting enterprises, while the manufacturing sector in another city accounts for 19.41% of the province's total electricity consumption among polluting enterprises. Other cities and counties account for 37.71%.
[0069] After lossy compression of the raw data, the polluting enterprises were categorized based on the data, as shown in Table 7:
[0070] Table 7 Classification of Polluting Enterprises
[0071] Industry type Non-destructive enterprise classification Culture, sports and entertainment industry First Tier Mining Second Tier Construction industry Second Tier Scientific research and technical services industry Second Tier Retail Second Tier Wholesale Second Tier manufacturing Third Tier
[0072] The two compression methods yielded no difference in the classification of polluted industries, as shown in Tables 6 and 7.
[0073] Finally, lossless compression and lossy compression methods were used to compress the acquired experimental data, and the compression results were recorded. The results after compression were compared in a comparative experiment.
[0074] The current data of a certain province were subjected to lossless and lossy compression respectively. Through experiments, the compressed data volume was obtained, as shown in Table 7.
[0075] Table 7 Experimental Results
[0076]
[0077] We can see that the lossy compression ratio is significantly higher than that of lossless compression. The lossy compression ratio is reduced by several times. Under the same application effect, lossy compression reduces the storage space by more than 107 times compared to lossless compression, which can save more than 107 times the cost.
[0078] In conclusion, lossy compression saves more storage space than lossless compression when the application effect is the same, making the proposed lossy compression scheme feasible. Lossy compression of IoT data is feasible, as evidenced by successful applications in video data. Video compression technology utilizes the characteristic that humans are insensitive to certain frequency components in images or sound waves, losing some information during compression. While the original video data cannot be completely recovered, the impact of the lost information on our eyes and hearing is negligible. However, lossy compression of video data can achieve a much higher compression ratio, hence the term "lossy compression, lossless application." The abundance of video software on the market today is not surprising. Successful cases of video data compression can be promoted, and practical examples show that it is also feasible in IoT applications; lossy compression and lossless application are operable in the IoT.
[0079] In the description of this invention, it should be understood that the terms "coaxial," "bottom," "one end," "top," "middle," "other end," "upper," "side," "top," "inner," "front," "center," "both ends," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0080] Furthermore, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first," "second," "third," or "fourth" may explicitly or implicitly include at least one of those features.
[0081] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "setting," "connection," "fixing," "screw connection," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal connection of two components or the interaction between two components. Unless otherwise explicitly limited, those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0082] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A lossless application of lossy compression for massive amounts of time-series data, characterized in that, Includes the following steps: S1. First, obtain the experimental data, and then organize and store it. S2. Then, the experimental data after lossy compression is decompressed, and the application value of the data after lossy compression and decompression is analyzed. S3. Next, load warning analysis is performed on the lossless compression experimental data and the experimental data after lossy compression, and the warning situation of the experimental data after lossy compression is compared with the warning situation of the original experimental data. S4. Then, the original experimental data is sorted according to the degree of contamination. At the same time, the original experimental data is lossily compressed, and the experimental data after lossy compression is sorted again. Then, the difference between the sorting of the original experimental data and the sorting of the experimental data after lossy compression is compared. S5. Finally, lossless compression and lossy compression methods were used to compress the obtained experimental data, and the compression results were recorded. The results after compression were compared in a comparative experiment. S6. Incorporate lossy compressed data into lossless compressed data; The experimental data in step S1 is current data; The data application value in step S2 refers to whether the data can achieve the expected purpose when actually used; The specific operation of the data application value analysis in step S2 is as follows: obtain the experimental data after lossy compression and decompression, and compare and analyze it with the experimental data after compression using lossless compression method to analyze whether the experimental data after lossy compression and decompression can achieve the same effect as the experimental data after compression using lossless compression method. The specific operation of performing load early warning analysis on the original experimental data in step S3 is as follows: First, analyze the load situation of the experimental data, set the early warning load, then conduct the experiment and record the number of early warnings and the early warning time period. The specific operation of performing load warning analysis on the experimental data after lossy compression in step S3 is as follows: the experimental data is compressed using a lossy compression method, and the load of the experimental data after lossy compression is analyzed. At the same time, a warning load is set, and the number of warnings and the warning time period are recorded. The specific classifications in step S4 are as follows: the first tier is for low-power-consuming and polluting industries, i.e., daily power consumption is less than 10,000 kWh; the second tier is for medium-power-consuming and polluting industries, i.e., daily power consumption is between 10,000 and 1 million kWh; and the third tier is for high-power-consuming and polluting industries, i.e., daily power consumption exceeds 1 million kWh.
2. The lossless application of lossy compression for massive time-series data as described in claim 1, characterized in that, The experimental results recorded in step S5 represent the data compression rate.
Citation Information
Patent Citations
KR1018292620000B1