Data compression device, data compression method, and data compression program

The data compression device addresses the challenge of preserving continuous value information in discrete time series data by using a specified period and difference threshold to distinguish noise from meaningful data, resulting in efficient data reduction and improved machine learning accuracy.

WO2025094290A1PCT designated stage expired Publication Date: 2025-05-08MITSUBISHI ELECTRIC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/039338
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing data compression technologies for discrete time series data fail to effectively eliminate noise while preserving information about continuous value periods, especially when the continuous value period is longer than a certain threshold.

Method used

A data compression device and method that determines whether a continuous value in time series data is noise or not based on a specified period and a difference threshold. If the continuous value period is less than the specified period and the difference with the previous value is within the threshold, it is deemed noise. Otherwise, it is preserved in the compressed data.

Benefits of technology

This approach effectively reduces data storage costs by eliminating noise while maintaining important information about continuous value periods, thereby improving the accuracy of machine learning models trained on compressed data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023039338_08052025_PF_FP_ABST
    Figure JP2023039338_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A data compression device (100) that compresses discrete time series data including a plurality of successive values that comprise data points indicating one or more successive identical values over time is provided with a noise determination unit (11). Treating each successive value included in the time series data as a target successive value, the noise determination unit (11) determines that the target successive value is noise when the duration of the target successive value is less than a designated period and the difference between the value of the target successive value and the value of the data point immediately temporally before the starting point of the target successive value is less than or equal to a difference threshold value, and determines that the target successive value is not noise when the duration of the target successive value is greater than or equal to the designated period and the difference between the value of the target successive value and the value of the data point immediately temporally before the starting point of the target successive value is less than or equal to the difference threshold value, and generates compressed data, which is data not indicating each successive value that was determined to be noise, is data indicating each successive value that was determined not to be noise, and is data resulting from compressing the time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Data compression device, data compression method, and data compression program

[0001] The present disclosure relates to a data compression device, a data compression method, and a data compression program.

[0002] There is a technique for compressing discrete time-series data that contains many continuous values. Patent Document 1 discloses an example of such a technique.

[0003] Japanese Patent Application Laid-Open No. 2001-165712

[0004] According to the technology disclosed in Patent Document 1, time series data is compressed according to a predetermined time, i.e., the time series data is compressed without considering the length of the period of consecutive identical values. Note that, in this technology, when a value change greater than the dead band width occurs, data corresponding to the value change is saved. Therefore, this technology has a problem in that when a value changes within the dead band width within a predetermined time range, even if the value change is not considered to be due to noise, the compressed time series data may not contain information indicating the value change. Because a huge amount of data is used in the learning process of machine learning, it is preferable to maximize the data compression rate. Furthermore, change points are important in machine learning. Furthermore, learning using data from which noise has been removed generally has little impact on accuracy. Therefore, when compressing discrete time series data containing many consecutive values ​​for the purpose of generating data to be used in the learning process of machine learning, it is preferable to remove noise while retaining information indicating the data during the period of consecutive identical values ​​if the period of consecutive identical values ​​is longer than a certain period, even if the value changes within the dead band width.

[0005] The present disclosure aims to provide a technology for compressing discrete time-series data containing many continuous values, which eliminates noise while retaining information indicating the data for a period in which the same value continues for a specified period or longer, even if the value changes within a range within a difference threshold.

[0006] A data compression device according to the present disclosure compresses discrete time series data including a plurality of continuous values ​​each consisting of data points showing one or more identical values ​​that are consecutive in time series, and is equipped with: a noise determination unit that, when each continuous value included in the time series data is a target continuous value, determines the target continuous value to be noise if the duration of the target continuous value is less than a specified period and a difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in time series is equal to or less than a difference threshold; determines the target continuous value to be not noise if the duration of the target continuous value is equal to or greater than the specified period and a difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in time series is equal to or less than the difference threshold; and generates compressed data, which is data that does not indicate each of the continuous values ​​determined to be noise but indicates each of the continuous values ​​determined to be not noise, and is data obtained by compressing the time series data.

[0007] According to the present disclosure, the noise determination unit determines a target continuous value to be noise if the duration of the target continuous value is less than a specified period and the difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than a difference threshold. The noise determination unit also determines a target continuous value to be non-noise if the duration of the target continuous value is equal to or greater than a specified period and the difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than a difference threshold. Furthermore, the noise determination unit generates compressed data, which is data obtained by compressing time-series data that does not represent each continuous value determined to be noise and represents each continuous value determined not to be noise. Here, the time-series data is discrete data including multiple continuous values. Therefore, according to the present disclosure, a technique for compressing discrete time-series data including many continuous values ​​can remove noise while retaining information indicating the data during a period in which the same value continues for a specified period or longer, even if the value changes within a range within a difference threshold.

[0008] FIG. 1 is a diagram showing an example of the configuration of a data compression system 90 according to a first embodiment. FIG. 2 is a diagram showing an example of the configuration of a data compression system 90 according to the first embodiment. FIG. 3 is a diagram explaining the processing of a noise determination unit 11 according to the first embodiment. FIG. 4 is a diagram explaining the processing of the noise determination unit 11 according to the first embodiment, where (a) is a diagram explaining a steady state and a transient state, and (b) is a diagram showing a function corresponding to the performance of an air conditioner. FIG. 4 is a diagram explaining the processing of the noise determination unit 11 according to the first embodiment. FIG. 5 is a diagram showing an example of the hardware configuration of a data compression device 100 according to the first embodiment. FIG. 6 is a flowchart showing the processing of the data compression system 90 according to the first embodiment. FIG. 7 is a flowchart showing the processing of the noise determination unit 11 according to the first embodiment. FIG. 8 is a diagram showing an example of the hardware configuration of a data compression device 100 according to a modified example of the first embodiment. FIG. 9 is a diagram showing an example of the configuration of a data compression system 90 according to a second embodiment. FIG. 10 is a diagram showing an example of the configuration of a statistic selection unit 30 according to the second embodiment. FIG. 11 is a diagram explaining a designated period 12 and a difference threshold 13 according to the second embodiment. FIG. 12 is a flowchart showing the processing of the data compression system 90 according to the second embodiment. FIG. 13 is a flowchart showing the processing of the statistic selection unit 30 according to the second embodiment. 1 is a flowchart showing the processing of a specified period calculation unit 32 according to embodiment 2. 2 is a flowchart showing the processing of a difference threshold calculation unit 34 according to embodiment 2. 3 is a diagram showing an example of the configuration of a data compression system 90 according to embodiment 3. 4 is a diagram showing an example of the configuration of a parameter determination unit 40 according to embodiment 3. 5 is a diagram showing an example of the configuration of a statistic selection unit 30 according to embodiment 3. 6 is a diagram explaining the processing of the data compression system 90 according to embodiment 3. 7 is a flowchart showing the processing of the data compression system 90 according to embodiment 3. 8 is a flowchart showing the processing of the data compression system 90 according to embodiment 3. 9 is a flowchart showing the processing of the statistic selection unit 30 according to embodiment 3. 10 is a flowchart showing the processing of the parameter determination unit 40 according to embodiment 3.

[0009] In the description of the embodiments and drawings, the same and corresponding elements are given the same reference numerals. The description of elements given the same reference numerals will be omitted or simplified as appropriate. Arrows in the drawings mainly indicate the flow of data or the flow of processing. Furthermore, "unit" or "device" may be read as "system," "circuit," "process," "procedure," "processing," or "circuitry" as appropriate.

[0010] Embodiment 1. This embodiment will be described in detail below with reference to the drawings. In this embodiment, a technique for compressing discrete time-series data containing many continuous values ​​with the aim of generating data to be used in the learning process of machine learning aims to simultaneously reduce the amount of data by removing noise that has no or only a small effect on machine learning, and to make the accuracy of a machine learning model trained based on data from which the noise has been removed equivalent to the accuracy of a machine learning model trained using raw data. Furthermore, this embodiment aims to reduce the amount of data by removing weak noise from the time-series data and compressing it in order to reduce the storage costs of the time-series data. The time-series data according to this embodiment is data consisting of discrete values ​​and including multiple continuous values. In other words, the time-series data is discrete time-series data. Continuous values ​​consist of data points that exhibit one or more identical values ​​that are consecutive in time series. As a specific example, continuous values ​​occur frequently in use cases of the Internet of Things (IoT). This can be attributed to factors such as low sensor sensitivity (resolution) or intentionally rounding values ​​by rounding off or the like to reduce the amount of data. An isolated point is sometimes referred to as a continuous value, where the value of the isolated point is different from the values ​​of any of the data points adjacent to the isolated point.

[0011] ***Description of Configuration*** Figures 1 and 2 show an example configuration of a data compression system 90 according to this embodiment. The data compression device 100 compresses time-series data. The data compression device 100 may be realized by a cloud system 1 as shown in Figure 1, or may be realized by an edge system 19 as shown in Figure 2. The data compression device 100 and the data restoration device 200 may be integrated into one unit. Below, a specific example in which the data compression device 100 is realized by the cloud system 1 will be described. The case in which the data compression device 100 is realized by the edge system 19 is the same as the case in which the data compression device 100 is realized by the cloud system 1. The restoration method 22 is data indicating a method for restoring the compressed data 14.

[0012] The time series data 5 acquired by each sensor 2 is transmitted to the cloud system 1 via the network 6. The time series data 5 is data in which the sensor values ​​acquired by each sensor 2 are recorded in time series. The time series data 5 corresponds to raw data. The cloud system 1 includes a data receiving unit 7, a data compression device 100, a database 15, and a data restoration device 200. The data receiving unit 7 receives the time series data 5. The data compression device 100 includes an encoding unit 9 and a noise determination unit 11.

[0013] The encoding unit 9 receives the time series data 5 from the data receiving unit 7 and generates encoded data 10 by encoding the received time series data 5. As a specific example, the encoding unit 9 generates the encoded data 10 by differential encoding. Differential encoding is a method of compressing discrete time series data based on the difference between adjacent data points. In differential encoding, each data point is retained if the difference between the data point and the immediately preceding data point in the time series is not zero, and is deleted if the difference between the data point and the immediately preceding data point in the time series is zero.

[0014] The noise determination unit 11 determines whether a partial data point sequence included in the encoded data 10 is noise based on a specified period 12 and a difference threshold 13, and deletes the partial data point sequence determined to be noise from the encoded data 10 to generate compressed data 14. The noise determination unit 11 may generate the compressed data 14 by differential encoding. The noise determination unit 11 stores the generated compressed data 14 in a database 15. The compressed data 14 is data that does not indicate each consecutive value determined to be noise by the noise determination unit 11, but indicates each consecutive value determined to be not noise by the noise determination unit 11, and is data obtained by compressing the time-series data 5. The compressed data 14 may also be data obtained by further compressing the encoded data 10. In this embodiment, the specified period 12 and the difference threshold 13 are introduced to determine whether each data point is noise. The specified period 12 and the difference threshold 13 may each be a value calculated based on domain knowledge. Note that the noise determination unit 11 may use the time-series data 5 instead of the encoded data 10. The compressed data 14 corresponds to data indicating features extracted from the raw data. The noise determination unit 11 determines a certain continuous value as noise if the duration of the certain continuous value is shorter than the specified period 12. However, even if the noise determination unit 11 determines that a certain continuous value is noise based on the specified period 12, the noise determination unit 11 does not determine that the certain continuous value is noise if the difference (fluctuation) between the value of the certain continuous value and the value of the data point immediately preceding the start point of the certain continuous value in the time series is greater than the difference threshold 13. When each continuous value included in the time series data 5 is a target continuous value, the noise determination unit 11 determines that the target continuous value is noise if the duration of the target continuous value is shorter than the specified period 12 and the difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than the difference threshold 13. Furthermore, the noise determination unit 11 determines that the target continuous value is not noise if the duration of the target continuous value is equal to or greater than the specified period 12 and the difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than the difference threshold 13.In addition, the noise determination unit 11 determines that the target continuous value is not noise if the difference between the value of the target continuous value and the value of the data point immediately preceding the starting point of the target continuous value in the time series is greater than the difference threshold 13.

[0015] FIG. 3 is a diagram illustrating a specific example of processing by the noise determination unit 11. In FIG. 3, the specified period 12 is 3, and the difference threshold 13 is 0.5. FIG. 3 shows original data and saved data. The original data is each data point indicated by the time-series data 5. The saved data is each data point indicated by the time-series data 5 that is saved in the compressed data 14 and is determined based on differential encoding. The duration of data D1 is 0, i.e., less than the specified period 12. Furthermore, the difference between the value of data D1 and the value of the continuous value S1 is 0.5, i.e., less than or equal to the difference threshold 13. Here, the value of the continuous value S1 is the value of the data point immediately preceding data D1 in the time series. Data D1 is also a continuous value. Therefore, the noise determination unit 11 determines data D1 to be noise and deletes data D1 from the encoded data 10. In the example shown in FIG. 3 , the value at time t1 in the time series data 5 and the encoded data 10 is 1.5, but the value at time t1 in the data restored from the compressed data 14 is 1. That is, in the compressed data 14, data D1 is essentially treated as data D1'. Here, the duration of each of the continuous values ​​S1 and S2 is 2. However, because data D1 is treated as data D1', the continuous values ​​S1, D1', and S2 are collectively treated as the continuous value S1'. Since the duration of the continuous value S1' is 6, i.e., equal to or greater than the specified period 12, the noise determination unit 11 does not determine the continuous value S1' as noise. Note that the noise determination unit 11 may also regard the entire continuous values ​​S1 and S2 as consecutive data points. The duration of the continuous value S3 is 3, i.e., equal to or greater than the specified period 12. Therefore, the noise determination unit 11 does not determine the continuous value S3 as noise. The duration of the continuous value S4 is 2. However, the difference between the value of continuous value S4 and the value of continuous value S3 is greater than the difference threshold 13. Therefore, the noise determination unit 11 does not determine continuous value S4 as noise. Similarly, the noise determination unit 11 does not determine continuous value S5 and continuous value S6 as noise.

[0016] The database 15 is a database for storing the generated compressed data 14 .

[0017] The data restoration device 200 includes a restoration unit 17. The restoration unit 17 acquires compressed data 14 from a database 15 and restores the acquired compressed data 14 to time-series data as restored data 18. In this case, the restoration unit 17 uses a decoding algorithm corresponding to the algorithm used to encode the compressed data 14.

[0018] An example of calculating the designated period 12 and the difference threshold 13 when the time-series data 5 is a room temperature data series of an air conditioner will be described using Figure 4. Here, the designated period 12 and the difference threshold 13 are each calculated based on domain knowledge. Note that the example shown in Figure 4 aims to detect abnormal conditions that are expected based on the performance of the air conditioner and the characteristics of the room in which the air conditioner is installed.

[0019] When the air conditioner is operating, the room temperature goes through a transient state and then to a steady state. The transient state is a state in which the room temperature is rising or falling toward the set temperature. The steady state is a state in which the room temperature remains stationary at the set temperature. After that, as shown in FIG. 4(a), after the air conditioning stops due to the steady state being reached, the room temperature gradually falls or rises over z minutes, and if the difference between the set temperature and the room temperature exceeds the threshold value, the air conditioner again transitions to a transient state. Here, at time x i+z and time x i The difference between the set temperature and the specified time period 12 is z minutes. Also, the specified time period 12 is 10 minutes. If the threshold value is α, the range of fluctuation within which the room temperature naturally drops or rises, i.e., without any environmental factors, can be defined as "|(set temperature) - (room temperature)| < α." If α is the room temperature at which the room temperature can be returned to the set temperature within the specified time period 12 based on the capacity of the air conditioner, then "α ≦ (difference threshold 13)" is obtained. If the room temperature value changes beyond the difference threshold 13 within the specified time period 12, it is considered that some environmental factor, such as ventilation or people entering or leaving the room, has occurred, and compressed data 14 is generated so that the changed room temperature value is preserved.

[0020] In setting the difference threshold 13, it is sufficient to know that "the air conditioner will raise or lower the room temperature by y degrees over x minutes" based on the capacity of the air conditioner (for example, 2.8 kW for cooling and 3.6 kW for heating) and the size of the room in which the air conditioner is installed. If the change in room temperature y can be defined as a function y = f(x) of the air conditioner's performance f and time x, the maximum fluctuation range per unit time can be determined as shown in Figure 4(b). In Figure 4(b), (x i+1 -x i ) corresponds to unit time, and "|f(x i ) -f(x i+1 ) | is the maximum fluctuation range of the room temperature that can be assumed based only on the performance of the air conditioner in a unit time. Here, if the granularity of the room temperature measured by the temperature sensor is 0.5 degrees, then time x i From time x i+1 The difference threshold 13 corresponding to is as shown in [Equation 1]. As a specific example, if the change in room temperature at a certain point in time is f(x 1 = 1) = 15.5, and f(x 2 = 2) = 18.5, the difference threshold 13 based on the domain knowledge is as shown in [Equation 2].

[0021]

[0022] A use case other than air conditioners includes data from a vibration sensor used to detect tool wear in machine tools installed in factories, etc., as shown in Figure 5. Here, the vibration waveform itself changes periodically and does not contain continuous values, as shown on the left side of Figure 5. However, when the vibration waveform is transformed using a Fourier transform or the like, data containing many continuous values ​​of a certain type may appear, as shown on the right side of Figure 5. The amount of data can be reduced by compressing such data by removing weak noise based on a specified period 12 and a difference threshold 13. Furthermore, information indicating normal and abnormal waveforms is retained in the compressed data 14.

[0023] 6 shows an example of the hardware configuration of a data compression device 100 according to this embodiment. The data compression device 100 is made up of a computer. The data compression device 100 may also be made up of multiple computers.

[0024] As shown in the figure, the data compression device 100 is a computer including hardware such as a processor 51, a memory 52, an auxiliary storage device 53, an input / output IF (Interface) 54, and a communication device 55. These pieces of hardware are connected as appropriate via signal lines 59.

[0025] The processor 51 is an integrated circuit (IC) that performs arithmetic processing and controls the hardware of a computer. Specific examples of the processor 51 include a central processing unit (CPU), a digital signal processor (DSP), or a graphics processing unit (GPU). The data compression device 100 may include multiple processors that replace the processor 51. The multiple processors share the role of the processor 51.

[0026] The memory 52 is typically a volatile storage device, and a specific example is RAM (Random Access Memory). The memory 52 is also called a primary storage device or a main memory. Data stored in the memory 52 is saved in the secondary storage device 53 as needed.

[0027] The auxiliary storage device 53 is typically a non-volatile storage device, and specific examples thereof include a ROM (Read Only Memory), an HDD (Hard Disk Drive), or a flash memory. Data stored in the auxiliary storage device 53 is loaded into the memory 52 as needed. The memory 52 and the auxiliary storage device 53 may be configured integrally.

[0028] The input / output IF 54 is a port to which an input device and an output device are connected. Specific examples of the input / output IF 54 include a USB (Universal Serial Bus) terminal. Specific examples of the input device include a keyboard and a mouse. Specific examples of the output device include a display.

[0029] The communication device 55 is a receiver and a transmitter, and is specifically a communication chip or a NIC (Network Interface Card).

[0030] Each unit of the data compression device 100 may use the input / output IF 54 and the communication device 55 as appropriate when communicating with other devices.

[0031] The auxiliary storage device 53 stores a data compression program. The data compression program causes a computer to realize the functions of each unit included in the data compression device 100. The data compression program is loaded into the memory 52 and executed by the processor 51. The functions of each unit included in the data compression device 100 are realized by software.

[0032] Data used when executing the data compression program and data obtained by executing the data compression program are stored in a storage device as appropriate. Each part of the data compression device 100 uses a storage device as appropriate. Specific examples of the storage device include at least one of the memory 52, the auxiliary storage device 53, a register in the processor 51, and a cache memory in the processor 51. Note that the terms "data" and "information" may have the same meaning. The storage device may be independent of the computer. The functions of the memory 52 and the auxiliary storage device 53 may be realized by other storage devices.

[0033] The data compression program may be recorded on a computer-readable non-volatile recording medium. Specific examples of the non-volatile recording medium include an optical disk and a flash memory. The data compression program may be provided as a program product.

[0034] ***Explanation of Operation*** The operation procedure of the data compression system 90 corresponds to a data compression method. Also, the program that realizes the operation of the data compression system 90 corresponds to a data compression program.

[0035] 7 is a flowchart showing an example of the processing performed by the data compression device 100. The processing performed by the data compression device 100 will be described with reference to FIG.

[0036] (Step S101 ) The data receiving unit 7 receives the time series data 5 and sends the received time series data 5 to the encoding unit 9 .

[0037] (Step S102) The encoding unit 9 receives the time series data 5 from the data receiving unit 7, encodes the received time series data 5 to generate encoded data 10, and sends the generated encoded data 10 to the noise determination unit 11.

[0038] (Step S103 ) The designated period 12 is input to the noise determination unit 11 .

[0039] (Step S104 ) The difference threshold 13 is input to the noise determination unit 11 .

[0040] (Step S105) The noise determination unit 11 receives the encoded data 10 and generates compressed data 14 using the received encoded data 10, the input specified period 12, and the difference threshold 13. At this time, the noise determination unit 11 detects consecutive values ​​whose duration is shorter than the specified period 12, and deletes from the encoded data 10 each consecutive value whose difference between the value of the consecutive value and the value of the data point immediately preceding the start point of the consecutive value in the time series is equal to or less than the difference threshold 13.

[0041] (Step S106) The noise determination unit 11 outputs the generated compressed data 14.

[0042] 8 and 9 are flowcharts showing an example of the process of the noise determination unit 11. The process of the noise determination unit 11 will be described with reference to FIGS.

[0043] (Step S121 ) The encoded data 10 is input to the noise determination unit 11 .

[0044] (Step S122 ) The designated period 12 is input to the noise determination unit 11 .

[0045] (Step S123 ) The difference threshold value 13 is input to the noise determination unit 11 .

[0046] (Step S124) The noise determination unit 11 sets the constant P to the value indicated by the specified period 12, the constant D to the value indicated by the difference threshold 13, and the encoded data 10 to (X, Y). Here, X[k] (1≦k≦n) indicates the kth time in the time series data 5, and Y[k] indicates the value at the kth time in the time series data 5. n is the maximum value of the subscripts of the time series data 5. E[k] (0≦E[k]≦n) indicates the subscript of X corresponding to the kth data point in the time series of the encoded data 10.

[0047] (Step S125) The noise determination unit 11 assigns 1 to the variable i and sets E' to an empty set. E' successively stores E's determined to be noise from the encoded data 10.

[0048] (Step S126) If i<n is satisfied, the noise determination unit 11 proceeds to step S127, otherwise the noise determination unit 11 proceeds to step S131.

[0049] (Step S127) If X[E[i+1]-1]-X[E[i]]<P is satisfied, the noise determination unit 11 proceeds to step S128. Otherwise, the noise determination unit 11 proceeds to step S130.

[0050] (Step S128) If |Y[E[i+1]-1]-Y[E[i]]|≦D is satisfied, the noise determination unit 11 proceeds to step S129. Otherwise, the noise determination unit 11 proceeds to step S130.

[0051] (Step S129) The noise determination unit 11 adds E[i] to E′.

[0052] (Step S130) The noise determination unit 11 increments the value of the variable i by one.

[0053] (Step S131) ​​The noise determination unit 11 determines E'' as the difference set between E and E', where A\B indicates the difference set obtained by subtracting set B from set A.

[0054] (Step S132) The noise determination unit 11 defines the compressed data 14 as (X[E''], Y[E'']).

[0055] (Step S133) The noise determination unit 11 outputs the generated compressed data 14.

[0056] ***Description of Effects of First Embodiment*** As described above, according to this embodiment, when time-series data contains many consecutive values, data points are not saved after a predetermined time has elapsed, as described in Patent Document 1, thereby enabling further compression of the time-series data. Furthermore, according to this embodiment, even if a fluctuation equal to or less than the difference threshold 13 occurs, if consecutive values ​​immediately following the fluctuation continue for a specified period 12 or longer, information indicating the consecutive values ​​is retained in the compressed data 14. Therefore, important change points in machine learning are not discarded. In machine learning, values ​​with large fluctuations, i.e., values ​​with large variance, often have a greater impact on predictions. This embodiment actively retains points with large fluctuations and removes only noise that has little or no effect on the accuracy of the machine learning model. Therefore, this embodiment reduces data volume and data storage costs while maintaining accuracy. To restore the compressed data 14, a typical differential encoding restoration algorithm is used as a specific example. Therefore, this embodiment does not require a special data restoration device. Note that the set values ​​of the specified period 12 and the difference threshold 13 do not affect data restoration, so the specified period 12 and the difference threshold 13 do not need to be retained as metadata. Furthermore, the process of generating the compressed data 14 may have the same effect as preprocessing in machine learning, so depending on the machine learning model (learning algorithm), improved accuracy can be expected.

[0057] ***Other Configurations*** <Modification 1> Fig. 10 shows an example of the hardware configuration of a data compression device 100 according to this modification. The data compression device 100 includes a processing circuit 58 instead of the processor 51, the processor 51 and memory 52, the processor 51 and auxiliary storage device 53, or the processor 51, memory 52, and auxiliary storage device 53. The processing circuit 58 is hardware that realizes at least a portion of the components included in the data compression device 100. The processing circuit 58 may be dedicated hardware, or may be a processor that executes a program stored in the memory 52.

[0058] When processing circuitry 58 is dedicated hardware, processing circuitry 58 may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. Data compression device 100 may include multiple processing circuits that replace processing circuit 58. The multiple processing circuits share the role of processing circuit 58.

[0059] In the data compression device 100, some of the functions may be realized by dedicated hardware, and the remaining functions may be realized by software or firmware.

[0060] The processing circuitry 58 is realized by, for example, hardware, software, firmware, or a combination of these. The processor 51, memory 52, auxiliary storage device 53, and processing circuitry 58 are collectively referred to as the "processing circuitry." In other words, the functions of the functional components of the data compression device 100 are realized by the processing circuitry. Data compression devices 100 according to other embodiments may also have a configuration similar to that of this modification.

[0061] Second Embodiment The following mainly describes the differences from the above-described embodiment with reference to the drawings.

[0062] ***Description of Configuration*** Fig. 11 shows an example of the configuration of a data compression system 90 according to this embodiment. The data compression device 100 according to this embodiment further includes a statistics selection unit 30, a specified period calculation unit 32, and a difference threshold calculation unit 34. Note that the data compression device 100 may be realized by an edge system 19.

[0063] 12 shows an example of the configuration of the statistic selection unit 30. The statistic selection unit 30 appropriately selects a statistic 31 and a statistic 33 from a statistic list 301. The statistic 31 is a statistic used when calculating the specified period 12. The statistic 33 is a statistic used when calculating the difference threshold 13. Note that the statistic 31 and the statistic 33 may be different from each other. The statistic list 301 is a list showing each statistic. The statistic list 301 is made up of data showing, as specific examples, the "arithmetic mean," the "median," the "mode," the "trimmed mean," and the "weighted mean."

[0064] The designated period calculation unit 32 calculates the designated period 12 based on the duration of each continuous value included in each time series data 5. Specifically, the designated period calculation unit 32 calculates the designated period 12 using the statistics 31, past time series data 35, and time series data 5 as input. The past time series data 35 is data restored from compressed data 14 generated in the past. The past time series data 35 may be time series data 5 acquired in the past. As a specific example, when the designated period 12 is determined by arithmetic averaging, the designated period calculation unit 32 calculates the designated period 12 using Equation 3. When the time series data 5 shown in FIG. 13 is input, the designated period calculation unit 32 calculates the designated period 12 as shown in Equation 4.

[0065]

[0066] The difference threshold calculation unit 34 calculates the difference threshold 13 based on the difference between the value of each continuous value included in each time series data 5 and the value of the data point immediately preceding the start point of each continuous value in the time series. Specifically, the difference threshold calculation unit 34 calculates the difference threshold 13 using the statistics 33, past time series data 35, and time series data 5 as input. As a specific example, when the difference threshold 13 is determined by arithmetic averaging, the difference threshold calculation unit 34 calculates the difference threshold 13 using Equation 5. When the time series data 5 shown in FIG. 13 is input, the specified period calculation unit 32 calculates the difference threshold 13 as shown in Equation 6.

[0067]

[0068] ***Explanation of Operation*** Fig. 14 is a flowchart showing an example of the processing of the data compression system 90. The processing of the data compression system 90 will be described with reference to Fig. 14 .

[0069] (Step S201 ) The statistic selection unit 30 selects the statistic 31 and the statistic 33 .

[0070] (Step S202) The time series data 5 and the past time series data 35 are input to the designated period calculation unit 32 and the difference threshold calculation unit 34, respectively.

[0071] (Step S203) The designated period calculation unit 32 calculates the designated period 12.

[0072] (Step S204) The difference threshold calculation unit 34 calculates the difference threshold 13.

[0073] 15 is a flowchart showing an example of the processing of the statistics selection unit 30. The processing of the statistics selection unit 30 will be described with reference to FIG.

[0074] (Step S221) The statistics list 301 is input to the statistics selection unit 30.

[0075] (Step S222 ) The statistics selection unit 30 selects, as the statistics 31 , statistics to be used when calculating the designated period 12 from the statistics indicated in the statistics list 301 .

[0076] (Step S223 ) The statistic selection unit 30 outputs the selected statistic 31 to the designated period calculation unit 32 .

[0077] (Step S224 ) The statistic selection unit 30 selects, as the statistic 33 , a statistic to be used when calculating the difference threshold 13 from the statistic listed in the statistic list 301 .

[0078] (Step S225) The statistic selection unit 30 outputs the selected statistic 33 to the difference threshold calculation unit .

[0079] 16 is a flowchart showing an example of the process of the designated period calculation unit 32. The process of the designated period calculation unit 32 will be described with reference to FIG.

[0080] (Step S241 ) The statistics 31 are input to the designated period calculation unit 32 .

[0081] (Step S242) The time-series data 5 is input to the designated period calculation unit 32.

[0082] (Step S243) The past time series data 35 is input to the designated period calculation unit 32.

[0083] (Step S244) The designated period calculation unit 32 assigns the time series data 5 to the variable Seq.

[0084] (Step S245) If the past time series data 35 exists, the designated period calculation unit 32 proceeds to step S246. Otherwise, the designated period calculation unit 32 proceeds to step S247.

[0085] (Step S246) The designated period calculation unit 32 adds the past time series data 35 to the variable Seq.

[0086] (Step S247) The designated period calculation unit 32 extracts n consecutive values ​​as period[i] (1≦i≦n) from the time-series data indicated by the variable Seq.

[0087] (Step S248) The designated period calculation unit 32 calculates the designated period 12 according to the input statistics 31.

[0088] (Step S249 ) The designated period calculation unit 32 outputs the calculated designated period 12 to the noise determination unit 11 .

[0089] 17 is a flowchart showing an example of the process of the difference threshold calculation unit 34. The process of the difference threshold calculation unit 34 will be described with reference to FIG.

[0090] (Step S261) The statistical quantity 33 is input to the difference threshold calculation unit .

[0091] (Step S262) The time-series data 5 is input to the difference threshold calculation unit 34.

[0092] (Step S263) The past time-series data 35 is input to the difference threshold calculation unit 34.

[0093] (Step S264) The difference threshold calculation unit 34 assigns the time-series data 5 to the variable Seq.

[0094] (Step S265) If the past time-series data 35 exists, the difference threshold calculation unit 34 proceeds to step S266. Otherwise, the difference threshold calculation unit 34 proceeds to step S267.

[0095] (Step S266) The difference threshold calculation unit 34 adds the past time-series data 35 to the variable Seq.

[0096] (Step S267) The difference threshold calculation unit 34 extracts the difference between m pieces of data from the time-series data indicated by the variable Seq as diff[i] (1≦i≦m).

[0097] (Step S268 ) The difference threshold calculation unit 34 calculates the difference threshold 13 according to the input statistics 33 .

[0098] (Step S269 ) The difference threshold calculation unit 34 outputs the calculated difference threshold 13 to the noise determination unit 11 .

[0099] ***Explanation of Effects of Embodiment 2*** As described above, according to this embodiment, the designated period 12 and the difference threshold 13 are calculated based on time-series data and selected statistics. Therefore, it is not necessary to determine the designated period 12 and the difference threshold 13 for each data series, i.e., no domain knowledge is required. Furthermore, according to this embodiment, the number of user parameters is reduced, improving ease of use.

[0100] Third Embodiment Hereinafter, differences from the above-described embodiments will be mainly described with reference to the drawings.

[0101] ***Description of Configuration*** Fig. 18 shows an example of the configuration of a data compression system 90 according to this embodiment. Compared to the data compression device 100 according to the second embodiment, the data compression device 100 further includes a parameter determination unit 40. Note that the data compression device 100 according to the first embodiment may also further include the parameter determination unit 40.

[0102] 19 shows an example configuration of the parameter determination unit 40. The parameter determination unit 40 includes a model generation unit 401, an inference unit 404, and an accuracy determination unit 407. The parameter determination unit 40 determines whether to adopt each of the calculated designated period 12 and the calculated difference threshold 13, based on the accuracy of the machine learning model trained using data restored from the compressed data 14 relative to the accuracy of the machine learning model trained using the time-series data 5.

[0103] The model generation unit 401 generates a comparison model 402 using the time-series data 5, and generates a determination target model 403 using the restored data 41. The restored data 41 is data restored from the compressed data 14 by the restoration unit 17, and is time-series data to be determined. Each of the comparison model 402 and the determination target model 403 is a machine learning model. The model generation unit 401 may use past time-series data 35 when generating the comparison model 402. The past time-series data 35 is time-series data generated by the restoration unit 17 retrieving past compressed data 14 from the database 15 and restoring the retrieved compressed data 14.

[0104] The inference unit 404 calculates an inference result 405 by performing inference using the comparison model 402, and calculates an inference result 406 by performing inference using the target model 403. Each of the inference result 405 and the inference result 406 is a predicted value. The inference result 406 corresponds to the target to be judged.

[0105] The accuracy determination unit 407 measures the inference error using the inference result 405 and the inference result 406, and determines whether the accuracy of the machine learning model corresponding to the compressed data 14 is acceptable.

[0106] FIG. 20 shows an example configuration of the statistic selection unit 30 according to this embodiment. The statistic selection unit 30 further includes a used table 302. The used table 302 is table data that holds each statistic determined by the parameter determination unit 40 and manages each used statistic. As shown in FIG. 20 , the used table 302 may hold a pair of a statistic corresponding to a specified period 12 and a statistic corresponding to a difference threshold 13. The used table 302 may include a used table corresponding to the specified period 12 and a used table corresponding to the difference threshold 13. The statistic selection unit 30 selects a statistic not included in the used table 302 from the statistic list 301. When a pair of statistics is held in the used table 302, the statistic selection unit 30 may select a combination of statistics not included in the used table 302. The statistic selection unit 30 records each used statistic in the used table 302.

[0107] FIG. 21 illustrates an example of processing performed by a data compression system 90 according to this embodiment. The parameter determination unit 40 creates machine learning models using the raw data and the time-series data restored from the compressed data 14, and determines whether the designated period 12 and the difference threshold 13 used to generate the compressed data 14 are acceptable based on the difference in inference accuracy between the two created machine learning models. If the parameter determination unit 40 rejects the designated period 12 and the difference threshold 13, the statistics selection unit 30 reselects the statistics used in calculating the designated period 12 and the difference threshold 13. The parameter determination unit 40 may determine whether to accept only one of the designated period 12 and the difference threshold 13, or may determine whether to accept the designated period 12 and the difference threshold 13 simultaneously or sequentially. Note that the machine learning model compared with the machine learning model corresponding to the restored data 41 may be a machine learning model trained using previously collected data. The parameter determination unit 40 may select a combination of the designated period 12 and the difference threshold 13 that minimizes the error relative to the accuracy of the machine learning model corresponding to the raw data.

[0108] ***Explanation of Operation*** Figures 22 and 23 are flowcharts showing an example of the processing of the data compression system 90. The processing of the data compression system 90 will be described with reference to Figures 22 and 23.

[0109] (Step S301) The restoration unit 17 restores the compressed data 14 generated by the noise determination unit 11 to generate restored data 41.

[0110] (Step S302) The time-series data 5 is input to the parameter determination unit 40.

[0111] (Step S303) The restoration unit 17 extracts each compressed data 14 generated in the past from the database 15, and restores each extracted compressed data 14 as past time-series data 35.

[0112] (Step S304) The parameter determination unit 40 measures the accuracy of the machine learning model corresponding to the restored data 41, and determines whether or not to adopt the current specified period 12 and difference threshold 13 based on the measurement result.

[0113] (Step S305) If the current specified period 12 and difference threshold 13 are to be adopted, the data compression system 90 proceeds to step S106. If they are not to be adopted, the data compression system 90 repeatedly executes the processing of this flowchart.

[0114] 24 is a flowchart showing an example of the processing of the statistics selection unit 30. The processing of the statistics selection unit 30 will be described with reference to FIG.

[0115] (Step S321) The used table 302 is input to the statistics selection unit 30.

[0116] (Step S322) If the selected statistics 31 and 33 exist in the used table 302, the statistics selection unit 30 returns to the process of selecting each statistic. Otherwise, the statistics selection unit 30 proceeds to step S323.

[0117] (Step S323 ) The statistic selection unit 30 records the selected statistic 31 and statistic 33 in the used table 302 .

[0118] 25 and 26 are flowcharts showing an example of the processing of the parameter determining unit 40. The processing of the parameter determining unit 40 will be described with reference to FIGS.

[0119] (Step S341) The time-series data 5 is input to the parameter determination unit 40.

[0120] (Step S342) The past time-series data 35 is input to the parameter determining unit 40.

[0121] (Step S343) The parameter determining unit 40 receives the restored data 41 to be determined.

[0122] (Step S344) The model generating unit 401 assigns the time-series data 5 to the variable BaseSeq.

[0123] (Step S345) If the past time series data 35 exists, the model generation unit 401 proceeds to step S346. Otherwise, the model generation unit 401 skips step S346.

[0124] (Step S346) The model generating unit 401 adds the past time-series data 35 to the variable BaseSeq.

[0125] (Step S347) The model generation unit 401 divides the data indicated by the variable BaseSeq into training data (TrainData) and evaluation data (TestData).

[0126] (Step S348) The model generation unit 401 generates a comparison model 402 by executing machine learning using the TrainData as input.

[0127] (Step S349) The inference unit 404 inputs the Test Data to the comparison model 402 generated by the model generation unit 401 and executes inference, thereby generating an inference result 405.

[0128] (Step S350) The model generation unit 401 divides the data indicated by the restored data 41 into training data (TrainData) and evaluation data (TestData).

[0129] (Step S351) The model generation unit 401 generates a target model 403 by executing machine learning using the TrainData as input.

[0130] (Step S352) The inference unit 404 inputs the TestData to the target model 403 generated by the model generation unit 401 and executes inference, thereby generating an inference result 406.

[0131] (Step S353) The accuracy determination unit 407 calculates the error of the inference result 406 relative to the inference result 405. Specific examples of the error index include RSME (Root Mean Squared Error), MSE (Mean Squared Error), MAE (Mean Absolute Error), or MAPE (Mean Absolute Percentage Error).

[0132] (Step S354) If the calculated error is equal to or greater than the predetermined threshold, the accuracy determining unit 407 proceeds to step S355. Otherwise, the accuracy determining unit 407 proceeds to step S356.

[0133] (Step S355) The accuracy determination unit 407 rejects the adoption / rejection 43.

[0134] (Step S356) The accuracy determination unit 407 determines the adoption / rejection 43 as adoption.

[0135] (Step S357) The parameter determination unit 40 outputs the adoption / rejection 43.

[0136] ***Explanation of Effects of Embodiment 3*** As described above, according to this embodiment, it is determined whether to adopt the designated period 12 and the difference threshold 13 based on the error between the machine learning model corresponding to the compressed data 14 and the machine learning model corresponding to the raw data. Therefore, according to this embodiment, it is possible to automatically determine the designated period 12 and the difference threshold 13 that minimize the degradation of the inference accuracy of the machine learning model.

[0137] ***Other Embodiments*** The above-described embodiments can be freely combined, or any of the components of each embodiment can be modified, or any of the components can be omitted from each embodiment. Furthermore, the embodiments are not limited to those shown in embodiments 1 to 3, and various modifications are possible as needed. The procedures described using flowcharts, etc., can be modified as appropriate.

[0138] 1 Cloud system, 2 Sensor, 5 Time series data, 6 Network, 7 Data receiving unit, 9 Encoding unit, 10 Encoded data, 11 Noise determination unit, 12 Specified period, 13 Difference threshold, 14 Compressed data, 15 Database, 17 Decompression unit, 18 Decompressed data, 19 Edge system, 22 Decompression method, 30 Statistics selection unit, 301 Statistics list, 302 Used table, 31 Statistics, 32 Specified period calculation unit, 33 Statistics, 34 Difference threshold calculation unit, 35 Past time series data, 40 Parameter determination unit, 401 Model generation unit, 402 Comparison model, 403 Determination target model, 404 Inference unit, 405, 406 Inference result, 407 Accuracy determination unit, 41 Decompressed data, 43 Acceptance / rejection, 51 Processor, 52 Memory, 53 Auxiliary storage device, 54 Input / output IF, 55 Communication device, 58 processing circuit, 59 signal line, 90 data compression system, 100 data compression device, 200 data recovery device.

Claims

1. A data compression device that compresses discrete time-series data including a plurality of continuous values ​​each consisting of one or more data points that are consecutive in time and have the same value, comprising: a noise determination unit that, when each continuous value included in the time-series data is a target continuous value, determines the target continuous value to be noise if the duration of the target continuous value is less than a specified period and a difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than a difference threshold; determines the target continuous value to be not noise if the duration of the target continuous value is equal to or greater than the specified period and a difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than the difference threshold; and generates compressed data, which is data that does not indicate each of the continuous values ​​determined to be noise, but indicates each of the continuous values ​​determined to be not noise, and is data obtained by compressing the time-series data.

2. The data compression device of claim 1, wherein the noise determination unit determines that the target continuous value is not noise if the difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is greater than the difference threshold.

3. The data compression device according to claim 1 or 2, wherein the noise determination section generates the compressed data by differential encoding.

4. The data compression device according to claim 1, wherein each of the specified period and the difference threshold value is a value calculated based on domain knowledge.

5. The data compression device according to any one of claims 1 to 3, further comprising: a designated period calculation unit that calculates the designated period based on the duration of each successive value included in the time series data; and a difference threshold calculation unit that calculates the difference threshold based on the difference between the value of each successive value included in the time series data and the value of a data point immediately preceding the start point of each successive value in the time series.

6. The data compression device according to claim 5, further comprising: a statistics selection unit that selects each of the statistics used when calculating the specified period and the statistics used when calculating the difference threshold.

7. The data compression device according to claim 5 or 6, further comprising a parameter determination unit that determines whether or not to adopt each of the calculated specified period and the calculated difference threshold value based on the accuracy of a machine learning model trained using data restored from the compressed data relative to the accuracy of a machine learning model trained using the time series data.

8. A data compression method executed by a data compression device which is a computer that compresses discrete time-series data including a plurality of continuous values ​​each consisting of one or more data points which are consecutive in time and which have the same value, wherein when each continuous value included in the time-series data is a target continuous value, the data compression device determines that the target continuous value is noise if the duration of the target continuous value is less than a designated period and a difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than a difference threshold, determines that the target continuous value is not noise if the duration of the target continuous value is equal to or more than the designated period and a difference between the value of the target continuous value and the value of the data point immediately preceding the start point of the target continuous value in the time series is equal to or less than the difference threshold, and generates compressed data which is data which does not indicate each of the continuous values ​​determined to be noise, but indicates each of the continuous values ​​determined to be not noise, and is data obtained by compressing the time-series data.

9. A data compression program executed by a data compression device, which is a computer that compresses discrete time-series data including a plurality of continuous values ​​each consisting of one or more data points that are consecutive in time and have the same value, the data compression program causing the data compression device to execute a noise determination process for generating compressed data, which is data obtained by compressing the time-series data and which is data that does not indicate each of the continuous values ​​determined to be noise, and which indicates each of the continuous values ​​determined to be not noise, and which is data obtained by compressing the time-series data, the data compression program performing a noise determination process for generating compressed data, which is data that does not indicate each of the continuous values ​​determined to be noise, and which is data that indicates each of the continuous values ​​determined to be not noise, when each continuous value included in the time-series data is a target continuous value, the data compression program performing a noise determination process for generating compressed data, which is data obtained by compressing the time-series data and which is data that does not indicate each of the continuous values ​​determined to be noise, and which indicates each of the continuous values ​​determined to be not noise.

Citation Information

Patent Citations

  • Method and device for storing data

    JP2001165712A

  • Compressor for time series observation data

    JP1991055919A

  • Data compression apparatus, data compression method, and program thereof

    JP2007104271A