Time-Series Compression Using Statistical Indices and Missing Values
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data compression techniques for time-series data lack the ability to determine the necessity and importance of individual data points, leading to reduced accuracy and isotropic compression, which is inappropriate for continuous data sets like battery deterioration or remaining capacity, where important information may be lost.
Innovation Solution
A time-series data processing method that quantizes and discretizes original data, calculates statistical indices, excludes data of low importance, performs compression, and stores both compressed data and statistical indices, allowing for later interpolation to restore original data quality, using machine learning for estimation and determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional data compression techniques are used to reduce data storage needs, then storage capacity is preserved, but data accuracy and important information are lost
Solution Approach 1:
The patent performs preliminary actions by calculating statistical indices (mean, standard deviation, normal distribution function) from the original time-series data before compression. These statistical characteristics are stored and later used during restoration to guide the interpolation process, ensuring that important data patterns are preserved while achieving compression.
Solution Approach 2:
The patent changes the representation parameters of the data by transforming raw time-series values into statistical indices (mean, standard deviation, normal distribution parameters). This parameter transformation allows the data to be compressed while retaining the essential characteristics needed for accurate restoration through inverse transformation and interpolation.
2Quantity of substance
If data is compressed to reduce storage needs, then storage capacity is preserved, but isotropic compression occurs which is inappropriate for continuous data sets
Solution Approach 1:
The patent applies local quality by treating different portions of the time-series data differently based on their statistical characteristics. By calculating and storing local statistical indices (mean, standard deviation, normal distribution function) for specific data segments, the compression adapts to the local properties of the data, preserving important variations while compressing less critical portions.
Solution Approach 2:
The patent introduces dynamics by making the compression process adaptive rather than static. The compression ratio and restoration accuracy are dynamically adjusted based on the statistical characteristics of the input data, allowing the system to optimize between compression efficiency and data fidelity for different types of time-series data.
3Measurement precision
If statistical indices are calculated and stored alongside compressed data, then data restoration accuracy is improved, but storage requirements increase
Solution Approach 1:
The patent creates simplified copies of the original data in the form of statistical indices (mean, standard deviation, normal distribution parameters) rather than storing the complete raw data. These statistical copies capture the essential characteristics of the time-series data, enabling accurate restoration while occupying minimal storage space compared to the original dataset.
Data Source
AI summary
A processing device for time-series data for reducing a data amount of time-series data, comprising: a compression processing unit that performs predetermined compression processing on original data of quantized time-series data, reduces the data amount, and converts the data into data for storage in a form that can be stored in a storage unit, wherein the compression processing unit calculates a first statistical index related to the original data, excludes data corresponding to an exclusion condition from the original data as a missing value, performs compression processing on the original data excluding a missing value, generates compressed data in which the data amount is reduced, and stores the compressed data and the first statistical index in the storage unit as data for storage.


