Data Compression Device, Data Compression Method, and Data Compression Program

The data compression device addresses the loss of important data points by using specified periods and difference thresholds to retain continuous values, improving compression efficiency and maintaining accuracy in machine learning applications.

JP7703130B1Active Publication Date: 2025-07-04MITSUBISHI ELECTRIC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025517517
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-07-04
Estimated Expiration
2043-10-31

AI Technical Summary

Technical Problem

Existing data compression techniques for discrete time-series data fail to retain information about continuous value periods longer than a certain length, leading to loss of important data points during compression, especially in machine learning applications where noise removal is crucial.

Method used

A data compression device that determines noise based on specified periods and difference thresholds, retaining continuous values longer than the specified period and with differences within the threshold, while deleting shorter continuous values with small differences.

Benefits of technology

This approach enhances data compression by retaining important change points, reducing data volume, and maintaining accuracy in machine learning models by removing noise without discarding relevant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007703130000007
    Figure 0007703130000007
  • Figure 0007703130000008
    Figure 0007703130000008
  • Figure 0007703130000009
    Figure 0007703130000009
Patent Text Reader

Abstract

A data compression device (100) that compresses discrete time series data including a plurality of continuous values composed of data points indicating one or more identical values that are serially continuous includes a noise determination unit (11). When each continuous value included in the time series data is a target continuous value, the noise determination unit (11) determines that the target continuous value is noise if the duration of the target continuous value is less than a specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to a difference threshold, and determines that the target continuous value is not noise if the duration of the target continuous value is greater than or equal to the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, and generates compressed data that is data not showing each continuous value determined to be noise, data showing each continuous value determined not to be noise, and data obtained by compressing the time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a data compression device, a data compression method, and a data compression program.

Background Art

[0002] There is a technique for compressing discrete time-series data that contains many continuous values. Patent Document 1 discloses an example of the technique.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] According to the technique disclosed in Patent Document 1, time-series data is compressed according to a predetermined time, that is, the time-series data is compressed without considering the length of the period during which the same value continues. In addition, when a change in a value greater than the deadband width occurs in the technique, data corresponding to the change in the value is stored. Therefore, in the technique, when the value changes within the deadband width within a predetermined time range, even when the change in the value is not considered to be caused by noise, the compressed time-series data may not include information indicating the change in the value. Here, since a huge amount of data is used in the learning process of machine learning, it is preferable to increase the compression rate of the data as much as possible. Also, change points are emphasized in machine learning. Furthermore, generally, the influence on accuracy is small even when learning is performed using data from which noise has been removed. Therefore, when compressing discrete time-series data containing many continuous values for the purpose of generating data used in the learning process of machine learning, while removing noise, even if the value changes within the deadband width, if the period during which the same value continues is longer than a certain length, it is preferable to leave information indicating the data during the period when the same value continues.

[0005] This disclosure aims to leave information indicating data during a period when the same value continues when the period during which the same value continues is longer than a specified period even if the value changes within the difference threshold range while removing noise in a technique for compressing discrete time-series data containing many continuous values.

Means for Solving the Problem

[0006] The data compression device according to this disclosure is a data compression device that compresses discrete time-series data including a plurality of continuous values composed of data points indicating one or more identical values that are continuous in time series, when each continuous value included in the time-series data is a target continuous value, if the continuous period of the target continuous value is less than the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, the target continuous value is determined to be noise, if the continuous period of the target continuous value is greater than or equal to the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, the target continuous value is determined not to be noise, By deleting each continuous value determined to be noise from the time-series data a noise determination unit that generates compressed data and is provided with.

Advantages of the Invention

[0007] According to the present disclosure, when the duration of a target continuous value is less than a specified period, and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to a difference threshold, the noise determination unit determines that the target continuous value is noise. Further, when the duration of the target continuous value is greater than or equal to the specified period, and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, the noise determination unit determines that the target continuous value is not noise. Furthermore, the noise determination unit generates compressed data that is data not showing each continuous value determined to be noise, data showing each continuous value determined not to be noise, and data obtained by compressing time series data. Here, the time series data is discrete data including a plurality of continuous values. Therefore, according to the present disclosure, in a technique for compressing discrete time series data including many continuous values, while excluding noise, information indicating data in a period in which the same value continues even when the value changes within the range within the difference threshold and the period in which the same value continues is greater than or equal to the specified period can be retained.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Embodiments for Carrying Out the Invention

[0009] In the description of the embodiments and the drawings, the same elements and corresponding elements are denoted by the same reference numerals. The description of the elements denoted by the same reference numerals may be omitted or simplified as appropriate. The arrows in the figures mainly indicate the flow of data or the flow of processing. Also, "section" or "device" may be appropriately read as "system", "circuit", "process", "procedure", "processing" or "circuitry".

[0010] Embodiment 1. Hereinafter, this embodiment will be described in detail with reference to the drawings. In this embodiment, in the technology of compressing discrete time-series data containing many continuous values for the purpose of generating data used in the learning process of machine learning, the data volume is reduced by removing noise that has no or little impact on machine learning, and the accuracy of a machine learning model trained based on the data with noise removed is made equivalent to the accuracy of a machine learning model trained using raw data. Also, in this embodiment, for the purpose of reducing the storage cost of time-series data, the data volume is reduced by removing weak noise from the time-series data and compressing it. The time-series data according to this embodiment is data consisting of discrete values and containing a plurality of continuous values. That is, the time-series data is discrete time-series data. The continuous values consist of data points indicating one or more identical values that are continuous in time series. As a specific example, in the use cases of IoT (Internet of Things), there are many continuous values. The factors include low sensitivity (resolution) of sensors, or intentionally rounding values by rounding off or the like to reduce the data volume. Note that outliers may also be expressed as continuous values. The value of an outlier is different from the values of any data points adjacent to the outlier.

[0011] ***Description of the configuration*** FIG. 1 and FIG. 2 show a configuration example of a data compression system 90 according to the present embodiment. The data compression device 100 compresses time-series data. The data compression device 100 may be realized by the cloud system 1 as shown in FIG. 1, or may be realized by the edge system 19 as shown in FIG. 2. The data compression device 100 and the data restoration device 200 may be integrally configured. Hereinafter, a specific example in which the data compression device 100 is realized by the cloud system 1 will be described. When the data compression device 100 is realized by the edge system 19, it is the same as the case where the data compression device 100 is realized by the cloud system 1. The restoration method 22 is data indicating a method for restoring the compressed data 14.

[0012] The time-series data 5 acquired by each sensor 2 is transmitted to the cloud system 1 via the network 6. The time-series data 5 is data in which the sensor values acquired by each sensor 2 are recorded in time series. The time-series data 5 corresponds to raw data. The cloud system 1 includes a data reception unit 7, a data compression device 100, a database 15, and a data restoration device 200. The data reception unit 7 receives the time-series data 5. The data compression device 100 includes an encoding unit 9 and a noise determination unit 11.

[0013] The encoding unit 9 receives the time-series data 5 from the data reception unit 7 and generates encoded data 10 by encoding the received time-series data 5. The symbolization unit 9 generates encoded data 10 by differential encoding as a specific example. Differential encoding is a method of compressing discrete time-series data based on the difference between adjacent data points. In differential encoding, each data point is retained when the difference from the immediately preceding data point in time series is not 0, and is deleted when the difference from the immediately preceding data point in time series is 0.

[0014] The noise determination unit 11 determines whether a partial data point sequence included in the encoded data 10 is noise based on the specified period 12 and the difference threshold 13, and deletes the partial data point sequence determined to be noise from the encoded data 10 to generate compressed data 14. The noise determination unit 11 may generate the compressed data 14 by differential encoding. The noise determination unit 11 stores the generated compressed data 14 in the database 15. The compressed data 14 is data that does not indicate each continuous value determined by the noise determination unit 11 to be noise, is data that indicates each continuous value determined by the noise determination unit 11 not to be noise, and is data obtained by compressing the time-series data 5. The compressed data 14 may be data obtained by further compressing the encoded data 10. In the present embodiment, the specified period 12 and the difference threshold 13 are introduced to determine whether each data point is noise. Each of the specified period 12 and the difference threshold 13 may be a value calculated based on domain knowledge. Note that the noise determination unit 11 may use the time-series data 5 instead of the encoded data 10. The compressed data 14 corresponds to data indicating features extracted from the raw data. When the continuous period of a certain continuous value is shorter than the specified period 12, the noise determination unit 11 determines that the certain continuous value is noise. However, even if the continuous value is determined to be noise based on the specified period 12, the noise determination unit 11 does not determine that the continuous value is noise when the difference (fluctuation) between the value of the continuous value and the value of the data point immediately preceding the start point of the continuous value in time series is greater than the difference threshold 13. When each continuous value included in the time-series data 5 is a target continuous value, the noise determination unit 11 determines that the target continuous value is noise when the duration of the target continuous value is less than the specified period 12 and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold 13. Further, the noise determination unit 11 determines that the target continuous value is not noise when the duration of the target continuous value is greater than or equal to the specified period 12 and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold 13. Note that the noise determination unit 11 determines that the target continuous value is not noise when the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is greater than the difference threshold 13.

[0015] FIG. 3 is a diagram for explaining a specific example of the process of the noise determination unit 11. In FIG. 3, the specified period 12 is 3 and the difference threshold 13 is 0.5. In FIG. 3, the original data and the stored data are shown. The original data is each data point indicated by the time-series data 5. The stored data is each data point stored in the compressed data 14 among the data points indicated by the time-series data 5, and is each data point determined based on differential encoding. The duration of the data D1 is 0, that is, less than the specified period 12. Further, the difference between the value of the data D1 and the value of the continuous value S1 is 0.5, that is, less than or equal to the difference threshold 13. Here, the value of the continuous value S1 is the value of the data point immediately before the data D1 in time series. The data D1 is also a continuous value. Therefore, the noise determination unit 11 determines that the data D1 is noise and deletes the data D1 from the encoded data 10. In the example shown in FIG. 3, although the value at time t1 in the time-series data 5 and the encoded data 10 is 1.5, the value at time t1 in the data obtained by restoring the compressed data 14 is 1. That is, in the compressed data 14, the data D1 is substantially treated as the data D1'. Here, the duration of each of the continuous values S1 and S2 is 2. However, since data D1 is treated as data D1', the continuous value S1, data D1', and the continuous value S2 are integrally treated as a continuous value S1'. The duration of the continuous value S1' is 6, that is, it is 12 or more of the specified period. Therefore, the noise determination unit 11 does not determine the continuous value S1' as noise. Note that the noise determination unit 11 may regard the entire continuous value S1 and the continuous value S2 as continuous data points. The duration of the continuous value S3 is 3, that is, it is 12 or more of the specified period. Therefore, the noise determination unit 11 does not determine the continuous value S3 as noise. The duration of the continuous value S4 is 2. However, the difference between the value of the continuous value S4 and the value of the continuous value S3 is greater than the difference threshold 13. Therefore, the noise determination unit 11 does not determine the continuous value S4 as noise. Similarly, the noise determination unit 11 does not determine each of the continuous values S5 and S6 as noise.

[0016] The database 15 is a database that stores the generated compressed data 14.

[0017] The data restoration device 200 includes a restoration unit 17. The restoration unit 17 acquires the compressed data 14 from the database 15 and restores the acquired compressed data 14 as restored data 18 into time-series data. At this time, the restoration unit 17 uses a decoding algorithm corresponding to the algorithm that encoded the compressed data 14.

[0018] An example of calculating the specified period 12 and the difference threshold 13 when the time-series data 5 is the room temperature data series of the air conditioner will be described with reference to FIG. 4. Here, each of the specified period 12 and the difference threshold 13 is calculated based on domain knowledge. Note that in the example shown in FIG. 4, the purpose is to detect a state that is not a normal state assumed from the performance of the air conditioner and the nature of the room in which the air conditioner is installed.

[0019] When the air conditioner operates, the room temperature reaches a steady state after going through a transient state. The transient state is a state where the room temperature is rising or falling towards the set temperature. The steady state is a state where the room temperature is stationary at the set temperature. After that, as shown in Fig. 4(a), after the air conditioner stops due to reaching the steady state, if the room temperature gradually drops or rises over z minutes and as a result, the difference between the set temperature and the room temperature exceeds the threshold value, the air conditioner transitions back to the transient state. Here, time x i+z and time x i The difference between them is z minutes. Also, assume the specified period 12 is 10 minutes. Let the threshold value be α, then the fluctuation range in which the room temperature naturally drops or rises without any environmental factors can be defined as "|(set temperature) - (room temperature)| < α". If α is set as the room temperature at which the air conditioner can return the room temperature to the set temperature within the specified period 12 based on its capacity, then "α ≤ (difference threshold 13)". If the value of the room temperature changes beyond the difference threshold 13 within the specified period 12, it is considered that some environmental factor such as ventilation or people entering and leaving has occurred, and the compressed data 14 is generated so that the changed value of the room temperature is saved.

[0020] Note that when setting the difference threshold 13, from the capacity of the air conditioner (specific examples: cooling 2.8kW, heating 3.6kW) and the size of the room where the air conditioner is installed, it is sufficient to know that "the air conditioner raises or lowers the room temperature by y degrees over x minutes". Here, if the change in room temperature y can be defined by the function y = f(x) of the performance f of the air conditioner and time x, the maximum fluctuation range per unit time can be specified as shown in Fig. 4(b). In Fig. 4(b), (x i+1 - x i ) corresponds to the unit time, and "|f(x i ) - f(x i+1 )|" is the maximum fluctuation range of the room temperature assumed only by the performance of the air conditioner per unit time. Here, when the granularity of the room temperature measured by the temperature sensor is in units of 0.5 degrees, from time x i to time x i+1The differential threshold value 13 corresponding thereto is as shown in [Equation 1]. As a specific example, when the change in room temperature at a certain point in time is f(x1 = 1) = 15.5 and f(x2 = 2) = 18.5, the differential threshold value 13 based on domain knowledge is as shown in [Equation 2].

[0021]

Equation

Equation

[0022] As a use case other than an air conditioner, data of a vibration sensor for detecting tool wear of a machine tool installed in a factory or the like as shown in FIG. 5 can be cited. Here, since the vibration waveform itself changes periodically, it does not include continuous values as shown on the left side of FIG. 5. However, when the vibration waveform is converted by Fourier transform or the like, data containing many certain continuous values may appear as shown on the right side of FIG. 5. By deleting and compressing weak noise based on the specified period 12 and the differential threshold value 13 from such data, the data volume can be reduced. Also, information indicating each of the normal waveform and the abnormal waveform is retained in the compressed data 14.

[0023] FIG. 6 shows a hardware configuration example of the data compression device 100 according to the present embodiment. The data compression device 100 consists of a computer. The data compression device 100 may consist of a plurality of computers.

[0024] The data compression device 100 is a computer including hardware such as a processor 51, a memory 52, an auxiliary storage device 53, an input / output IF (Interface) 54, and a communication device 55, as shown in this figure. These hardware components are appropriately connected via a signal line 59.

[0025] The processor 51 is an IC (Integrated Circuit) that performs arithmetic processing and controls the hardware of a computer. As a specific example, the processor 51 is a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The data compression device 100 may include a plurality of processors that replace the processor 51. The plurality of processors share the role of the processor 51.

[0026] The memory 52 is typically a volatile storage device and, as a specific example, is a RAM (Random Access Memory). The memory 52 is also referred to as the main storage device or main memory. The data stored in the memory 52 is saved in the auxiliary storage device 53 as needed.

[0027] The auxiliary storage device 53 is typically a non-volatile storage device and, as specific examples, is a ROM (Read Only Memory), an HDD (Hard Disk Drive), or a flash memory. The data stored in the auxiliary storage device 53 is loaded into the memory 52 as needed. The memory 52 and the auxiliary storage device 53 may be integrally configured.

[0028] The input / output IF 54 is a port to which an input device and an output device are connected. As a specific example, the input / output IF 54 is a USB (Universal Serial Bus) terminal. As specific examples, the input device is a keyboard and a mouse. As a specific example, the output device is a display.

[0029] The communication device 55 is a receiver and a transmitter. As a specific example, the communication device 55 is a communication chip or a NIC (Network Interface Card).

[0030] When each part of the data compression device 100 communicates with other devices or the like, the input / output IF 54 and the communication device 55 may be appropriately used.

[0031] The auxiliary storage device 53 stores a data compression program. The data compression program is a program that causes a computer to realize the functions of each part included in the data compression device 100. The data compression program is loaded into the memory 52 and executed by the processor 51. The functions of each part included in the data compression device 100 are realized by software.

[0032] Data used when executing the data compression program, data obtained by executing the data compression program, and the like are appropriately stored in a storage device. Each part of the data compression device 100 appropriately uses the storage device. The storage device includes, as specific examples, at least one of the memory 52, the auxiliary storage device 53, a register in the processor 51, and a cache memory in the processor 51. Note that the term "data" may have the same meaning as the term "information". The storage device may be independent of the computer. The functions of the memory 52 and the auxiliary storage device 53 may be realized by other storage devices.

[0033] The data compression program may be recorded on a computer-readable non-volatile recording medium. Specific examples of the non-volatile recording medium are an optical disk or a flash memory. The data compression program may be provided as a program product.

[0034] ***Description of Operations*** The operation procedure of the data compression system 90 corresponds to a data compression method. Also, a program for realizing the operation of the data compression system 90 corresponds to a data compression program.

[0035] FIG. 7 is a flowchart showing an example of the processing of the data compression device 100. The processing of the data compression device 100 will be described with reference to FIG. 7.

[0036] (Step S101) The data receiving unit 7 receives the time-series data 5 and sends the received time-series data 5 to the encoding unit 9.

[0037] (Step S102) The encoding unit 9 receives the time-series data 5 from the data receiving unit 7, generates the encoded data 10 by encoding the received time-series data 5, and sends the generated encoded data 10 to the noise determination unit 11.

[0038] (Step S103) The specified period 12 is input to the noise determination unit 11.

[0039] (Step S104) The difference threshold 13 is input to the noise determination unit 11.

[0040] (Step S105) The noise determination unit 11 receives the encoded data 10 and generates the compressed data 14 using the received encoded data 10, the input specified period 12, and the difference threshold 13. At this time, the noise determination unit 11 detects continuous values whose continuous period is shorter than the specified period 12, and deletes from the encoded data 10 each continuous value among the detected continuous values where the difference between the value of the continuous value and the value of the data point immediately before the start point of the continuous value in time series is less than or equal to the difference threshold 13.

[0041] (Step S106) The noise determination unit 11 outputs the generated compressed data 14.

[0042] FIG. 8 and FIG. 9 are flowcharts showing an example of the processing of the noise determination unit 11. The processing of the noise determination unit 11 will be described with reference to FIGS. 8 and 9.

[0043] (Step S121) The encoded data 10 is input to the noise determination unit 11.

[0044] (Step S122) The specified period 12 is input to the noise determination unit 11.

[0045] (Step S123) The differential threshold value 13 is input to the noise determination unit 11.

[0046] (Step S124) The noise determination unit 11 sets the constant P as the value indicated by the specified period 12, the constant D as the value indicated by the differential threshold value 13, and the encoded data 10 as (X, Y). Here, X[k] (1 ≤ k ≤ n) indicates the k-th time in time series of the time series data 5, and Y[k] indicates the value at the k-th time in time series of the time series data 5. n is the maximum value of the subscript of the time series data 5. E[k] (0 ≤ E[k] ≤ n) indicates the subscript of X corresponding to the k-th data point in time series of the encoded data 10.

[0047] (Step S125) The noise determination unit 11 substitutes 1 into the variable i and sets E' as an empty set. E that is determined to be noise among the encoded data 10 is sequentially stored in E'.

[0048] (Step S126) If i < n is satisfied, the noise determination unit 11 proceeds to step S127. Otherwise, the noise determination unit 11 proceeds to step S131.

[0049] (Step S127) If X[E[i + 1] - 1] - X[E[i]] < P is satisfied, the noise determination unit 11 proceeds to step S128. Otherwise, the noise determination unit 11 proceeds to step S130.

[0050] (Step S128) If |Y[E[i + 1] - 1] - Y[E[i]]| ≤ D is satisfied, the noise determination unit 11 proceeds to step S129. Otherwise, the noise determination unit 11 proceeds to step S130.

[0051] (Step S129) The noise determination unit 11 adds E[i] to E’.

[0052] (Step S130) The noise determination unit 11 increments the value of the variable i by 1.

[0053] (Step S131) The noise determination unit 11 sets E’’ as the difference set between E and E’. Here, A\B represents the difference set obtained by subtracting set B from set A.

[0054] (Step S132) The noise determination unit 11 sets the compressed data 14 to (X[E’’], Y[E’’]).

[0055] (Step S133) The noise determination unit 11 outputs the generated compressed data 14.

[0056] ***Description of the effects of Embodiment 1*** As described above, according to this embodiment, when the time-series data contains many continuous values, each data point at the time when a predetermined time has elapsed as shown in Patent Document 1 is not stored, so that the time-series data can be more compressed. Also, according to this embodiment, even when there is a variation of 13 or less in the difference threshold, if the continuous value immediately after the variation continues for 12 or more in the specified period, the information indicating the continuous value in the compressed data 14 is retained. Therefore, important change points in machine learning are not discarded. Here, in machine learning, values with large variations, that is, values with large variances, often have a greater impact on prediction. This embodiment actively retains points with large variations and deletes only noise that has no or little impact on the accuracy of the machine learning model. Therefore, according to this embodiment, while maintaining the accuracy, the data volume can be reduced and the data storage cost can be reduced. In the restoration of the compressed data 14, as a specific example, a normal differential encoding restoration algorithm is used. Therefore, in this embodiment, it is not necessary to prepare a special data restoration device. Note that since the set values of the specified period 12 and the difference threshold 13 do not affect the restoration of the data, the specified period 12 and the difference threshold 13 may not be left as metadata. Furthermore, the process of generating the compressed data 14 may sometimes have the same effect as preprocessing in machine learning. Therefore, depending on the machine learning model (learning algorithm), an improvement in accuracy is expected.

[0057] ***Other configurations*** <Modification Example 1> FIG. 10 shows a hardware configuration example of the data compression device 100 according to this modification example. Instead of the processor 51, the processor 51 and the memory 52, the processor 51 and the auxiliary storage device 53, or the processor 51, the memory 52, and the auxiliary storage device 53, the data compression device 100 includes a processing circuit 58. The processing circuit 58 is hardware that realizes at least a part of each unit included in the data compression device 100. The processing circuit 58 may be dedicated hardware, or may be a processor that executes a program stored in the memory 52.

[0058] When the processing circuit 58 is dedicated hardware, as a specific example, the processing circuit 58 is a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. The data compression device 100 may include a plurality of processing circuits that replace the processing circuit 58. The plurality of processing circuits share the role of the processing circuit 58.

[0059] In the data compression device 100, some functions may be realized by dedicated hardware, and the remaining functions may be realized by software or firmware.

[0060] The processing circuit 58 is realized by, for example, hardware, software, firmware, or a combination thereof. The processor 51, the memory 52, the auxiliary storage device 53, and the processing circuit 58 are collectively referred to as the "processing circuitry". That is, the functions of the respective functional components of the data compression device 100 are realized by the processing circuitry. The data compression device 100 according to another embodiment may have the same configuration as this modification example.

[0061] Embodiment 2. Hereinafter, mainly the differences from the above-described embodiment will be described with reference to the drawings.

[0062] ***Description of Configuration*** FIG. 11 shows a configuration example of the data compression system 90 according to this embodiment. The data compression device 100 according to this embodiment further includes a statistic selection unit 30, a specified period calculation unit 32, and a difference threshold calculation unit 34. Note that the data compression device 100 may be realized by the edge system 19.

[0063] FIG. 12 shows a configuration example of the statistic selection unit 30. The statistic selection unit 30 appropriately selects a statistic 31 and a statistic 33 from among the statistic list 301. The statistic 31 is a statistic used when calculating the specified period 12. The statistic 33 is a statistic used when calculating the difference threshold 13. Note that the statistic 31 and the statistic 33 may be different statistics from each other. The statistic list 301 is a list showing each statistic. The statistic list 301 is composed of data showing, for example, each of "arithmetic mean", "median", "mode", "trimmed mean", and "weighted mean".

[0064] The specified period calculation unit 32 calculates the specified period 12 based on the duration of each continuous value included in each time series data 5. Specifically, the specified period calculation unit 32 calculates the specified period 12 with the statistic 31, the past time series data 35, and the time series data 5 as inputs. The past time series data 35 is data obtained by restoring the compressed data 14 generated in the past. The past time series data 35 may be the time series data 5 acquired in the past. As a specific example, when determining the specified period 12 by the arithmetic mean, the specified period calculation unit 32 calculates the specified period 12 according to [Equation 3]. When the time series data 5 shown in FIG. 13 is input, the specified period calculation unit 32 calculates the specified period 12 as shown in [Equation 4].

[0065]

Equation

Equation

[0066] The difference threshold calculation unit 34 calculates the difference threshold 13 based on the difference between the value of each continuous value included in each time series data 5 and the value of the data point immediately before the start point of each continuous value in time series. Specifically, the difference threshold calculation unit 34 calculates the difference threshold 13 with the statistic 33, the past time series data 35, and the time series data 5 as inputs. As a specific example, when determining the difference threshold 13 by the arithmetic mean, the difference threshold calculation unit 34 calculates the difference threshold 13 according to [Equation 5]. When the time series data 5 shown in FIG. 13 is input, the specified period calculation unit 32 calculates the difference threshold 13 as shown in [Equation 6].

[0067]

Equation

Equation

[0068] ***Explanation of the operation*** Figure 14 is a flowchart showing an example of the processing of the data compression system 90. The processing of the data compression system 90 will be described with reference to Figure 14.

[0069] (Step S201) The statistic selection unit 30 selects the statistic 31 and the statistic 33.

[0070] (Step S202) The time series data 5 and the past time series data 35 are input to each of the specified period calculation unit 32 and the difference threshold calculation unit 34.

[0071] (Step S203) The specified period calculation unit 32 calculates the specified period 12.

[0072] (Step S204) The difference threshold calculation unit 34 calculates the difference threshold 13.

[0073] Figure 15 is a flowchart showing an example of the processing of the statistic selection unit 30. The processing of the statistic selection unit 30 will be described with reference to Figure 15.

[0074] (Step S221) The statistic list 301 is input to the statistic selection unit 30.

[0075] (Step S222) The statistic selection unit 30 selects, as the statistic 31, the statistic used when calculating the specified period 12 from among the statistics indicated by the statistic list 301.

[0076] (Step S223) The statistic selection unit 30 outputs the selected statistic 31 to the specified period calculation unit 32.

[0077] (Step S224) The statistic selection unit 30 selects, as the statistic 33, the statistic used when calculating the difference threshold 13 from among the statistics indicated by the statistic list 301.

[0078] (Step S225) The statistic selection unit 30 outputs the selected statistic 33 to the difference threshold calculation unit 34.

[0079] FIG. 16 is a flowchart showing an example of the process of the specified period calculation unit 32. The process of the specified period calculation unit 32 will be described with reference to FIG. 16.

[0080] (Step S241) The statistic 31 is input to the specified period calculation unit 32.

[0081] (Step S242) The time series data 5 is input to the specified period calculation unit 32.

[0082] (Step S243) The past time series data 35 is input to the specified period calculation unit 32.

[0083] (Step S244) The specified period calculation unit 32 substitutes the time series data 5 into the variable Seq.

[0084] (Step S245) If the past time series data 35 exists, the specified period calculation unit 32 proceeds to step S246. Otherwise, the specified period calculation unit 32 proceeds to step S247.

[0085] (Step S246) The specified period calculation unit 32 adds the past time series data 35 to the variable Seq.

[0086] (Step S247) The specified period calculation unit 32 extracts n consecutive values from the time series data indicated by the variable Seq as period[i] (1 ≤ i ≤ n).

[0087] (Step S248) The specified period calculation unit 32 calculates the specified period 12 according to the input statistic 31.

[0088] (Step S249) The specified period calculation unit 32 outputs the calculated specified period 12 to the noise determination unit 11.

[0089] FIG. 17 is a flowchart showing an example of the processing of the difference threshold calculation unit 34. The processing of the difference threshold calculation unit 34 will be described with reference to FIG. 17.

[0090] (Step S261) The statistical quantity 33 is input to the difference threshold calculation unit 34.

[0091] (Step S262) The time series data 5 is input to the difference threshold calculation unit 34.

[0092] (Step S263) The past time series data 35 is input to the difference threshold calculation unit 34.

[0093] (Step S264) The difference threshold calculation unit 34 substitutes the time series data 5 into the variable Seq.

[0094] (Step S265) If the past time series data 35 exists, the difference threshold calculation unit 34 proceeds to step S266. Otherwise, the difference threshold calculation unit 34 proceeds to step S267.

[0095] (Step S266) The difference threshold calculation unit 34 adds the past time series data 35 to the variable Seq.

[0096] (Step S267) The difference threshold calculation unit 34 extracts the differences of m data from the time series data indicated by the variable Seq as diff[i] (1 ≤ i ≤ m).

[0097] (Step S268) The difference threshold calculation unit 34 calculates the difference threshold 13 according to the input statistical quantity 33.

[0098] (Step S269) The difference threshold calculation unit 34 outputs the calculated difference threshold 13 to the noise determination unit 11.

[0099] ***Explanation of the effects of Embodiment 2*** As described above, according to the present embodiment, the specified period 12 and the difference threshold 13 are calculated based on the time series data and the selected statistic. Therefore, it is not necessary to determine each of the specified period 12 and the difference threshold 13 for each data series, that is, domain knowledge is not required. In addition, according to the present embodiment, since the user parameters are reduced, the usability is improved.

[0100] Embodiment 3. Hereinafter, mainly the differences from the above-described embodiments will be described with reference to the drawings.

[0101] ***Explanation of the configuration*** FIG. 18 shows a configuration example of a data compression system 90 according to the present embodiment. The data compression device 100 further includes a parameter determination unit 40 as compared with the data compression device 100 according to Embodiment 2. Note that the data compression device 100 according to Embodiment 1 may further include the parameter determination unit 40.

[0102] FIG. 19 shows a configuration example of the parameter determination unit 40. The parameter determination unit 40 includes a model generation unit 401, an inference unit 404, and an accuracy determination unit 407. The parameter determination unit 40 determines whether to adopt each of the calculated specified period 12 and the calculated difference threshold 13 based on the accuracy of the machine learning model learned using the compressed data 14 restored data with respect to the accuracy of the machine learning model learned using the time series data 5.

[0103] The model generation unit 401 generates a comparison model 402 using the time-series data 5, and generates a model to be determined 403 using the restored data 41. The restored data 41 is the data restored by the restoration unit 17 from the compressed data 14, and is the time-series data to be determined. Each of the comparison model 402 and the model to be determined 403 is a machine learning model. The model generation unit 401 may use the past time-series data 35 when generating the comparison model 402. The past time-series data 35 is time-series data generated by the restoration unit 17 retrieving past compressed data 14 from the database 15 and restoring the retrieved compressed data 14.

[0104] The inference unit 404 calculates an inference result 405 by performing inference using the comparison model 402, and calculates an inference result 406 by performing inference using the model to be determined 403. Each of the inference result 405 and the inference result 406 is a predicted value. The inference result 406 is the object to be determined.

[0105] The accuracy determination unit 407 determines whether the accuracy of the machine learning model corresponding to the compressed data 14 can be tolerated by measuring the inference error using the inference result 405 and the inference result 406.

[0106] FIG. 20 shows a configuration example of the statistic selection unit 30 according to the present embodiment. The statistic selection unit 30 further includes a used table 302. The used table 302 is table data that holds each statistic determined by the parameter determination unit 40, and is table data that manages each used statistic. As shown in FIG. 20, the used table 302 may hold a pair of a statistic corresponding to the specified period 12 and a statistic corresponding to the difference threshold 13. There may be a used table corresponding to the specified period 12 and a used table corresponding to the difference threshold 13 as the used table 302. The statistic selection unit 30 selects statistics not included in the used table 302 from the statistic list 301. When a set of statistics is held in the used table 302, the statistic selection unit 30 may select a combination of statistics not included in the used table 302. The statistic selection unit 30 records each used statistic in the used table 302.

[0107] FIG. 21 shows an image of the processing of the data compression system 90 according to the present embodiment. The parameter determination unit 40 creates a machine learning model using each of the raw data and the time series data obtained by restoring the compressed data 14, and determines whether to adopt the specified period 12 and the difference threshold 13 used when generating the compressed data 14 based on the difference in the inference accuracy between the two created machine learning models. When the parameter determination unit 40 does not adopt the specified period 12 and the difference threshold 13, the statistic selection unit 30 reselects the statistics used in the calculation of the specified period 12 and the difference threshold 13. The parameter determination unit 40 may execute the determination of whether to adopt only one of the specified period 12 and the difference threshold 13, or may execute the determination of whether to adopt the specified period 12 and the difference threshold 13 simultaneously or in order. Note that the machine learning model to be compared with the machine learning model corresponding to the restored data 41 may be a machine learning model learned using the data collected so far. The parameter determination unit 40 may adopt a combination of the specified period 12 and the difference threshold 13 for which the error with respect to the accuracy of the machine learning model corresponding to the raw data is minimized.

[0108] ***Explanation of Operations*** FIGS. 22 and 23 are flowcharts showing an example of the processing of the data compression system 90. The processing of the data compression system 90 will be described with reference to FIGS. 22 and 23.

[0109] (Step S301) The restoration unit 17 generates restored data 41 by restoring the compressed data 14 generated by the noise determination unit 11.

[0110] (Step S302) The time series data 5 is input to the parameter determination unit 40.

[0111] (Step S303) The restoration unit 17 retrieves each of the compressed data 14 generated in the past from the database 15, and restores the retrieved compressed data 14 as the past time series data 35.

[0112] (Step S304) The parameter determination unit 40 measures the accuracy of the machine learning model corresponding to the restored data 41, and determines whether to adopt the current specified period 12 and the difference threshold 13 based on the measurement result.

[0113] (Step S305) When adopting the current specified period 12 and the difference threshold 13, the data compression system 90 proceeds to step S106. When not adopting these, the data compression system 90 repeatedly executes the processing of this flowchart.

[0114] FIG. 24 is a flowchart showing an example of the processing of the statistic selection unit 30. The processing of the statistic selection unit 30 will be described with reference to FIG. 24.

[0115] (Step S321) The used table 302 is input to the statistic selection unit 30.

[0116] (Step S322) If the selected statistics 31 and 33 exist in the used table 302, the statistic selection unit 30 returns to the process of selecting each statistic. Otherwise, the statistic selection unit 30 proceeds to step S323.

[0117] (Step S323) The statistic selection unit 30 records the selected statistics 31 and 33 in the used table 302.

[0118] FIG. 25 and FIG. 26 are flowcharts showing an example of the processing of the parameter determination unit 40. The processing of the parameter determination unit 40 will be described with reference to FIGS. 25 and 26.

[0119] (Step S341) The time-series data 5 is input to the parameter determination unit 40.

[0120] (Step S342) The past time-series data 35 is input to the parameter determination unit 40.

[0121] (Step S343) The restoration data 41 to be determined is input to the parameter determination unit 40.

[0122] (Step S344) The model generation unit 401 substitutes the time-series data 5 into the variable BaseSeq.

[0123] (Step S345) If the past time-series data 35 exists, the model generation unit 401 proceeds to step S346. Otherwise, the model generation unit 401 skips step S346.

[0124] (Step S346) The model generation unit 401 adds the past time-series data 35 to the variable BaseSeq.

[0125] (Step S347) The model generation unit 401 divides the data indicated by the variable BaseSeq into learning data (TrainData) and evaluation data (TestData).

[0126] (Step S348) The model generation unit 401 generates a comparison model 402 by performing machine learning with TrainData as input.

[0127] (Step S349) The inference unit 404 generates an inference result 405 by inputting TestData to the comparison model 402 generated by the model generation unit 401 and performing inference.

[0128] (Step S350) The model generation unit 401 divides the data indicated by the restored data 41 into training data (TrainData) and evaluation data (TestData).

[0129] (Step S351) The model generation unit 401 generates a determination target model 403 by performing machine learning with TrainData as input.

[0130] (Step S352) The inference unit 404 generates an inference result 406 by inputting TestData to the determination target model 403 generated by the model generation unit 401 and performing inference.

[0131] (Step S353) The accuracy determination unit 407 calculates the error between the inference result 406 and the inference result 405. As specific examples of the error index, there are RSME (Root Mean Squared Error), MSE (Mean Squared Error), MAE (Mean Absolute Error), or MAPE (Mean Absolute Percentage Error).

[0132] (Step S354) If the calculated error is equal to or greater than a predetermined threshold, the accuracy determination unit 407 proceeds to step S355. Otherwise, the accuracy determination unit 407 proceeds to step S356.

[0133] (Step S355) The accuracy determination unit 407 sets the acceptance / rejection 43 to non-acceptance.

[0134] (Step S356) The accuracy determination unit 407 sets the acceptance / rejection 43 to acceptance.

[0135] (Step S357) The parameter determination unit 40 outputs an acceptance / rejection 43.

[0136] ***Explanation of the effects of Embodiment 3*** As described above, according to the present embodiment, based on the error of the machine learning model corresponding to the compressed data 14 with respect to the machine learning model corresponding to the raw data, the acceptance / rejection of the specified period 12 and the difference threshold 13 is determined. Therefore, according to the present embodiment, it is possible to automatically determine the specified period 12 and the difference threshold 13 that minimize the decrease in the inference accuracy of the machine learning model.

[0137] ***Other embodiments*** Free combinations of the above-described embodiments, or modifications of any components of each embodiment, or omissions of any components in each embodiment are possible. Also, the embodiments are not limited to those shown in Embodiments 1 to 3, and various changes can be made as necessary. The procedures described using flowcharts and the like may be changed as appropriate.

Explanation of reference numerals

[0138] 1 Cloud system, 2 Sensor, 5 Time-series data, 6 Network, 7 Data receiving unit, 9 Encoding unit, 10 Encoded data, 11 Noise determination unit, 12 Specified period, 13 Difference threshold, 14 Compressed data, 15 Database, 17 Restoration unit, 18 Restored data, 19 Edge system, 22 Restoration method, 30 Statistic selection unit, 301 Statistic list, 302 Used table, 31 Statistic, 32 Specified period calculation unit, 33 Statistic, 34 Difference threshold calculation unit, 35 Past time-series data, 40 Parameter determination unit, 401 Model generation unit, 402 Comparison model, 403 Model to be determined, 404 Inference unit, 405, 406 Inference results, 407 Accuracy determination unit, 41 Restored data, 43 Acceptance / rejection, 51 Processor, 52 Memory, 53 Auxiliary storage device, 54 Input / output IF, 55 Communication device, 58 Processing circuit, 59 Signal line, 90 Data compression system, 100 Data compression device, 200 Data restoration device.

Claims

1. A data compression device for compressing discrete time-series data including a plurality of continuous values composed of data points indicating one or more identical values continuously in a time series, when each continuous value included in the time-series data is a target continuous value, if the duration of the target continuous value is less than a specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in the time series is less than or equal to a difference threshold, the target continuous value is determined to be noise, if the duration of the target continuous value is greater than or equal to the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in the time series is less than or equal to the difference threshold, the target continuous value is determined not to be noise, a noise determination unit that generates compressed data by deleting each continuous value determined to be noise from the time-series data A data compression device comprising.

2. The data compression device according to claim 1, wherein the noise determination unit determines that the target continuous value is not noise when the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in the time series is greater than the difference threshold.

3. The data compression device according to claim 1 or 2, wherein the noise determination unit generates the compressed data by differential encoding.

4. The data compression device according to claim 1 or 2, wherein each of the specified period and the difference threshold is a value calculated based on domain knowledge.

5. The data compression device further includes, a specified period calculation unit that calculates the specified period based on the duration of each continuous value included in the time-series data, a difference threshold calculation unit that calculates the difference threshold based on the difference between the value of each continuous value included in the time-series data and the value of the data point immediately before the start point of each continuous value in the time series The data compression device according to claim 1 or 2, comprising.

6. The data compression device further includes, a statistic selection unit that selects each of the statistic used when calculating the specified period and the statistic used when calculating the difference threshold The data compression device according to claim 5, comprising.

7. The data compression device further includes, A parameter determination unit that determines whether to adopt each of the calculated specified period and the calculated difference threshold based on the accuracy of a machine learning model learned using the compressed data restored data with respect to the accuracy of a machine learning model learned using the time series data The data compression device according to claim 5, comprising the same

8. A data compression method executed by a data compression device that is a computer that compresses discrete time series data including a plurality of continuous values composed of data points indicating one or more identical values that are continuous in time series, When each continuous value included in the time series data is a target continuous value, The data compression device, When the duration of the target continuous value is less than the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, determine that the target continuous value is noise, When the duration of the target continuous value is equal to or longer than the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, determine that the target continuous value is not noise, A data compression method for generating compressed data by deleting each continuous value determined to be noise from the time series data

9. A data compression program executed by a data compression device that is a computer that compresses discrete time series data including a plurality of continuous values composed of data points indicating one or more identical values that are continuous in time series, When each continuous value included in the time series data is a target continuous value, When the duration of the target continuous value is less than the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, determine that the target continuous value is noise, When the duration of the target continuous value is equal to or longer than the specified period and the difference between the value of the target continuous value and the value of the data point immediately before the start point of the target continuous value in time series is less than or equal to the difference threshold, determine that the target continuous value is not noise, Noise determination processing for generating compressed data by deleting each continuous value determined to be noise from the time series data A data compression program that causes the data compression device to execute the same

Citation Information

Patent Citations

  • Compressor for time series observation data

    JP1991055919A

  • Data compression apparatus, data compression method, and program thereof

    JP2007104271A

  • Method and device for storing data

    JP2001165712A