Method and apparatus for processing time-series data

The method addresses the inefficiencies in compressing time-series data by calculating statistical indices and excluding less important data, resulting in efficient data reduction and preservation of critical information for machine learning.

JP7694522B2Active Publication Date: 2025-06-18TOYOTA JIDOSHA KK
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2022153454
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-06-18
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

Conventional data compression techniques struggle to determine the necessity and importance of each individual piece of data, particularly for time-series data that changes continuously over time, leading to inefficient compression and loss of important information.

Method used

A method for processing time-series data that involves quantizing or discretizing the data, calculating first statistical indices, excluding data meeting a predetermined condition as missing values, performing compression on the reduced data, and storing the compressed data with statistical indices. The method also includes a restoration process to generate learning data by interpolating missing values based on statistical indices.

Benefits of technology

This approach efficiently reduces the data volume while preserving important information, allowing for effective storage and restoration of time-series data, which is critical for machine learning applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694522000001
    Figure 0007694522000001
  • Figure 0007694522000002
    Figure 0007694522000002
  • Figure 0007694522000003
    Figure 0007694522000003
Patent Text Reader

Abstract

To provide a time series data processing method and processing apparatus capable of appropriately and efficiently reducing and a data volume and storing data while leaving important information.SOLUTION: The time series data processing apparatus reduces a data volume of time series data, and comprises a compression processing unit which performs prescribed compression processing on original data of quantized time series data to reduce the data volume into data for storage in such a form that it can be stored in a storage unit. The compression processing unit calculates a first statistical index about the original data (step S104), excludes data satisfying an exclusion condition from the original data as missing values (step S105), performs compression processing on the original data from which the missing values are excluded, to generate compressed data with the reduced data volume (step S106), and stores the compressed data and the first statistical index in the storage unit as the data for storage (step S107).SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technique for handling a large amount of data stored in a predetermined storage device or storage medium, and more particularly, to a technique for performing compression processing to reduce the data volume of time-series data that changes continuously over time and then using it.

Background Art

[0002] In recent years, the practical application of techniques for efficiently performing learning (or inference, recognition, judgment, etc.) processes by utilizing advanced technologies such as artificial intelligence (AI) and information and communication technology (ICT) has been promoted. Among them, machine learning is one of the learning methods executed by AI. Machine learning enables a machine (computer) to learn on its own using a large number of given data, and based on the learning result (trained model), it optimizes the output data for the input data. When storing such a large amount of data used as input data for machine learning in a predetermined storage device or storage medium, usually, data compression (or data volume reduction) processing is performed. An invention related to such data compression processing is described in Patent Document 1.

[0003] Patent Document 1 describes a method for compressing time-series data that decimates data based on a threshold value from within a numerical data sequence (time-series data). The method for compressing time-series data described in this Patent Document 1 uses a relationship between an expected value of a compression rate corresponding to a predetermined number of numerical data and a specific compression algorithm used for compression and a threshold value to perform a first step of initially setting the threshold value corresponding to a target compression rate, a second step of compressing the numerical data by a compression algorithm using the set threshold value and calculating an actual compression rate of the numerical data for each predetermined number, a third step of determining whether it is necessary to reset the threshold value based on the calculated actual compression rate and the target compression rate, and a fourth step of resetting the threshold value corresponding to the target compression rate using the most recent predetermined number of numerical data and the relationship between the expected value of the compression rate and the threshold value. Thereby, in the method for compressing time-series data described in this Patent Document 1, the threshold value serving as a reference when decimating and compressing time-series data is automatically set so that the deviation from the target compression rate is reduced.

[0004] Note that Patent Document 2 describes a learning data processing apparatus aimed at improving the quality of learning data. The learning data processing apparatus described in this Patent Document 2 includes a data processing unit that generates learning data used in a learning apparatus that generates a learning model based on time-series data including at least one type of measurement value. The data processing unit calculates at least one of a statistical value of measurement values included in one or a plurality of predetermined periods in the time-series data, or an outlier determination upper limit value or an outlier determination lower limit value based on the statistical value, and excludes from the time-series data measurement values that are at least one of greater than or equal to the outlier determination upper limit value or less than or equal to the outlier determination lower limit value among the measurement values included in the predetermined period. Alternatively, measurement values that satisfy a predetermined condition among the measurement values included in the time-series data are excluded from the time-series data.

[0005] In addition, Patent Document 3 describes a device for complementing time-series data of instantaneous heartbeats, which aims to enable appropriate spectral analysis even for instantaneous heartbeat data with a missing section caused by measurement anomalies or the like. The device for complementing time-series data of instantaneous heartbeats described in this Patent Document 3 calculates a complement value for a missing section of time-series data of instantaneous heartbeats having a missing section and the time of the missing section, and when the calculated time of the missing section is equal to or longer than the complement time and the time obtained by subtracting the complement time from the missing section is equal to or longer than the time to be complemented, the calculated complement value is complemented to the missing section. Further, in this Patent Document 3, it is disclosed that, for example, the average value of instantaneous heartbeats in an analysis target section is used as the above complement value.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0007] A large amount of data such as that used in the above-described machine learning is stored, for example, in an auxiliary storage device or an external storage medium (storage medium) of a server or a computer. At that time, in order to avoid the storage capacity of the storage device or the storage medium becoming strained, the data is stored in the storage device or the storage medium or the like in a state where the data amount has been reduced by performing predetermined processing. For example, in the method for compressing time-series data described in Patent Document 1 above, the data is thinned out under predetermined conditions, so that the data is irreversibly compressed and the data amount is reduced. Further, for example, an irreversible data compression format such as JPEG [Joint Photgraphic Experts Group] is widely used.

[0008] In such conventional data compression techniques, it is not easy to determine the necessity of each individual piece of data when compressing the data. For example, data related to image display data, character information, etc. is, relatively speaking, a collection of random or irregular individual data. When compressing such data, it is not possible to easily determine the necessity or importance of each individual piece of data. Therefore, for example, in the compression process of image display data by JPEG mentioned above, overall and on average, compression with reduced accuracy is performed, and an irreversible compression process with isotropy is carried out. On the other hand, for example, changes in the deterioration state, remaining capacity, or remaining life of a battery are continuous changes over time, and data related to the deterioration state and life of the battery becomes time-series data that changes continuously over time. Time-series data related to such continuous events does not become random or irregular data like image display data or character information, but basically becomes continuous data with a certain tendency. Therefore, when compressing time-series data related to such continuous events, it is desirable to determine the necessity and importance of each individual piece of data and compress those data, and compression with isotropy and reduced average accuracy as described above is not necessarily appropriate.

[0009] As described above, when storing, saving, or using time-series data that changes continuously over time, there is still room for improvement in order to appropriately and effectively compress the time-series data (reduce the data volume of the time-series data to be stored).

[0010] This invention was conceived by focusing on the above-mentioned technical problems, and aims to provide a processing method and a processing device for time-series data that can consider the necessity and importance of each individual piece of data in a large amount of time-series data, appropriately and efficiently reduce the data volume while leaving important information for storage.

Means for Solving the Problems

[0011] In order to achieve the above object, the present invention provides a method for processing time-series data that reduces the data volume of a large number of time-series data that change continuously over time. The method includes a compression processing step of performing a predetermined compression process on the original data of the time-series data that has been quantized or discretized, reducing the data volume to obtain storage data in a form that can be stored in a predetermined storage unit, and a restoration processing step of performing a restoration process on the storage data to obtain learning data in a form that can be used in a predetermined arithmetic process. The compression processing step includes a step of calculating a first statistical index (or basic statistic) related to the original data, a step of excluding, as missing values, data that meet a predetermined first condition for discriminating data with low importance for the arithmetic process from the original data, a step of performing the compression process on the original data with the missing values excluded to generate compressed data with the reduced data volume, and a step of storing the compressed data and the first statistical index in the storage unit as the storage data. The restoration processing step includes a step of reading the storage data from the storage unit, a step of generating restored data by performing a predetermined process for restoring the storage data on the storage data, a step of calculating a second statistical index related to the restored data, a step of specifying, as interpolation target locations, locations corresponding to data that meet a predetermined second condition from among the missing locations of the restored data including the traces of the excluded missing values, and a step of calculating an interpolation value to be applied to the interpolation target locations based on the first statistical index and the second statistical index, and supplementing (interpolating) the interpolation target locations with the interpolation values to generate the learning data approximating the original data from the restored data.

[0012] In addition, the data that meets the first condition in the method for processing time-series data of the present invention is data in the original data that is less than or equal to a lower threshold value set as a lower limit value and greater than or equal to an upper threshold value set as an upper limit value. The data that meets the second condition is data within a predetermined range including the lower threshold value or within a predetermined range including the upper threshold value for the data before and after the missing value in the time-series direction of the original data. The interpolated value may be extracted from a random number obtained based on the normal distribution function of the restored data among the second statistical indicators.

[0013] In addition, the arithmetic processing in the method for processing time-series data of the present invention is machine learning that performs learning based on a large number of input data and makes predictions or judgments based on the results of the learning. The learning data may be the input data in the machine learning.

[0014] In addition, the original data in the method for processing time-series data of the present invention is data related to a power storage device mounted on a vehicle, and the machine learning may be for predicting the change over time of the power storage device.

[0017] Furthermore, the present invention is a time-series data processing apparatus that reduces the data volume of a large number of time-series data that change continuously over time. From among the original data of the quantized time-series data, the original data in which data corresponding to a predetermined exclusion condition is excluded as a missing value is subjected to a predetermined compression process to reduce the data volume, and a restoration process is performed on the stored data stored in a predetermined storage device or storage medium to obtain learning data in a form that can be used in a predetermined arithmetic process. The restoration process unit includes: reading the stored data from the storage device or the storage medium; generating restored data obtained by performing a predetermined process for restoring the stored data on the stored data; calculating a second statistical index related to the restored data; identifying, as an interpolation target location, a location corresponding to data that meets a predetermined interpolation condition from among the missing locations of the restored data including a trace from which the missing value has been excluded; calculating an interpolation value to be applied to the interpolation target location based on a first statistical index related to the original data and the second statistical index; and generating the learning data approximating the original data by filling (interpolating) the interpolation target location with the interpolation value. This may be a feature of the present invention.

[0018] In addition, the data corresponding to the exclusion condition in the time-series data processing apparatus of the present invention is data in the original data that is less than or equal to a lower threshold value set as a lower limit value and greater than or equal to an upper threshold value set as an upper limit value. The data corresponding to the interpolation condition is data within a predetermined range including the lower threshold value or within a predetermined range including the upper threshold value for the data before and after the missing value in the time-series direction of the original data. The interpolation value may be extracted from a random number obtained based on the normal distribution function of the restored data among the second statistical indices.

[0019] And the arithmetic processing in the time-series data processing device of the present invention is machine learning that performs learning based on a large number of input data and makes predictions or judgments based on the results of the learning. The learning data is the input data in the machine learning, the original data is data related to a power storage device mounted on a vehicle, and the machine learning may predict the change over time of the power storage device.

Advantages of the Invention

[0020] The present invention, for example, performs a predetermined compression process on time-series data (original data) to reduce the data volume in order to efficiently store a large amount of time-series data used as input data (learning data) for machine learning in a storage unit. Further, a restoration process is performed on the time-series data (stored data) stored in the storage unit to restore it to a state where it can be appropriately used as learning data.

[0021] When compressing the original data, for example, first statistical indicators related to the original data, such as the average, standard deviation, and normal distribution function, are calculated. At the same time, the importance of each data is determined according to the first condition, and data with low importance is excluded as missing values. After reducing the missing values in this way, the compression process is performed. Then, the compressed data after the compression process is stored in the storage unit as stored data together with the first statistical indicators related to the original data. Therefore, it is possible to leave data with high importance, that is, to leave the characteristics of the original data, and efficiently and effectively reduce and compress the data volume of the original data and store it in the storage unit.

[0022] On the other hand, when restoring the stored data (compressed data) that has been compressed and stored, the first statistical index obtained from the original data and the second statistical index obtained from the restored data are used, and learning data approximated to the original data is generated. Specifically, the missing values (interpolation target locations) in the stored data are filled with interpolation values so that the shape of the normal distribution of the original data and the shape of the normal distribution of the learning data to be restored are approximated. That is, the missing values are approximately interpolated. Note that "interpolation" in this case is a process of calculating an approximate value for filling in the missing values, and is, for example, an arithmetic process similar to the mathematical "linear interpolation" method. Therefore, in the description of this invention, terms such as "interpolation" and "interpolate" are used instead of "complement" and "supplement".

[0023] Therefore, according to this invention, it is possible to efficiently reduce the data volume while considering the necessity and importance of each individual data in a large amount of time-series data and leaving important information. Therefore, it is possible to appropriately store a large amount of time-series data without straining the storage capacity of the storage unit. And the time-series data compressed and stored can be restored to closely approximate the original data and can be appropriately used as learning data.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0025] Embodiments of the present invention will be described with reference to the drawings. Note that the embodiments shown below are merely examples of the case where the present invention is embodied and do not limit the present invention.

[0026] The method and apparatus for processing time-series data according to embodiments of the present invention perform a predetermined compression process on a large amount of time-series data (original data) used as input data or learning data for machine learning using AI, for example, to reduce the data volume and save it. Then, a restoration process is performed on the saved time-series data (data for storage) to restore it to a state where it can be appropriately used as learning data. As an example, the apparatus for processing time-series data according to embodiments of the present invention is used when estimating the state of change over time (deterioration state) of a power storage device mounted on a vehicle by using machine learning with an existing general vehicle as the control target, and executes a compression process and a restoration process on the time-series data (original data). For this purpose, the apparatus for processing time-series data according to embodiments of the present invention includes a control unit mounted on the vehicle and a server installed outside the vehicle.

[0027] FIG. 1 shows an example of a vehicle equipped with a control unit as a component of the apparatus for processing time-series data according to an embodiment of the present invention. The vehicle Ve shown in FIG. 1 mainly includes a power source (POWER) 1 for driving, drive wheels 2, a starter motor 3, a battery 4, a detection unit 5, a control unit (ECU) 6, and a communication module (DCM) 7.

[0028] The power source 1 is a power source that outputs a driving torque for driving the vehicle Ve. The power source 1 is, for example, an internal combustion engine such as a gasoline engine or a diesel engine, and is configured such that the output adjustment and the operating states such as starting and stopping are electrically controlled. In the case of a gasoline engine, the opening degree of the throttle valve, the supply amount or injection amount of fuel, the execution and stop of ignition, and the ignition timing are electrically controlled. Alternatively, in the case of a diesel engine, the injection amount of fuel, the injection timing of fuel, or the opening degree of the throttle valve (in the EGR system) is electrically controlled. In the example shown in FIG. 1, an engine 8 equipped with a starter motor 3 described later is mounted as the power source 1.

[0029] Note that the driving force source 1 may be, for example, an electric motor such as a permanent magnet synchronous motor or an induction motor. In that case, the electric motor may be, for example, a so-called motor-generator that has both a function as a prime mover that is driven by the supply of electric power to output torque and a function as a generator that generates electricity by being driven by an external torque. If it is a motor-generator, the rotational speed, torque, or the switching between the function as a prime mover and the function as a generator is electrically controlled. Further, the driving force source 1 may be a so-called hybrid drive unit equipped with both the engine 8 and an electric motor (motor-generator).

[0030] The drive wheels 2 generate the driving force of the vehicle Ve when the driving torque output by the driving force source 1 is transmitted thereto. In the embodiment shown in FIG. 1, the drive wheels 2 are connected to the driving force source 1 via the transmission 9, the differential gear 10, and the drive shaft 11. Note that the vehicle Ve in the embodiment of this invention may be a front-wheel drive vehicle that transmits the driving torque to the front wheels and generates the driving force at the front wheels as in the embodiment shown in FIG. 1. Alternatively, the vehicle Ve may be a rear-wheel drive vehicle that transmits the driving torque to the rear wheels via, for example, a propeller shaft (not shown) or the like and generates the driving force at the rear wheels. Alternatively, the vehicle Ve may be a four-wheel drive vehicle provided with a transfer mechanism (not shown) that transmits the driving torque to both the front wheels and the rear wheels and generates the driving force at both the front wheels and the rear wheels.

[0031] The starter motor 3 is mounted on the engine 8 when the engine 8 is mounted as the driving force source 1 of the vehicle Ve as described above, and drives the crankshaft (not shown) when the engine 8 is started. The starter motor 3 operates when electric power is supplied from the battery 4 described later. Note that instead of the starter motor 3, an alternator (not shown) may function as the starter of the engine 8. Alternatively, a motor (not shown) having both the function of a starter and the function of an alternator may be used.

[0032] The battery 4 corresponds to the "power storage device" in the embodiment of the present invention and supplies power to the above-described starter motor 3. In the example shown in FIG. 1, the battery 4 is a so-called auxiliary battery and supplies power to in-vehicle devices (not shown) such as the lighting lamp and air conditioner of the vehicle Ve. Note that the "power storage device" in the embodiment of the present invention can also be targeted at, for example, the main battery or drive battery (not shown) in a hybrid vehicle or an electric vehicle.

[0033] The detection unit 5 is a device or apparatus for acquiring various data and information necessary when controlling the vehicle Ve, and includes, for example, a power supply unit, a microcomputer, sensors, and input / output interfaces. In particular, the detection unit 5 in the embodiment of the present invention has a battery voltage sensor 5a that detects the voltage of the battery 4 as information having a high degree of influence or relevance on the deterioration of the battery 4 when estimating the deterioration state of the "power storage device", that is, the battery 4. In addition, the detection unit 5 has a travel distance sensor 5b that detects the travel distance of the vehicle Ve, an outside air temperature sensor 5c that detects the outside air temperature, a battery temperature sensor 5d that detects the temperature of the battery 4, a battery resistance sensor 5e that detects the internal resistance of the battery 4, and a counter 5f that measures the number of starts of the engine 8, etc., as information that has a lower degree of influence or relevance on the deterioration of the battery 4 than the voltage of the battery 4 but that indirectly or comprehensively affects the deterioration of the battery 4. In addition, the detection unit 5 has, for example, a vehicle speed sensor (or wheel speed sensor) 5g that detects the vehicle speed and a rotation speed sensor 5h that detects the rotation speed of the engine 8. Further, the detection unit 5 may include, for example, a GPS [Global Positioning System] receiver (not shown) for acquiring the position information of the vehicle Ve and an in-vehicle camera (not shown) for acquiring imaging information regarding the external situation of the vehicle Ve. Then, the detection unit 5 is electrically connected to a control unit 6 described later and outputs an electrical signal corresponding to the detection value, calculation value, position information, etc. of the various sensors, devices, apparatuses, etc. as described above to the control unit 6 as detection data.

[0034] The control unit 6 is an electronic control device mainly composed of, for example, a microcomputer, and comprehensively controls the vehicle Ve. Various data detected or measured by the above detection unit 5 are input to the control unit 6. For this purpose, the control unit 6 has a data acquisition unit 6a described later. Then, the control unit 6 transmits various data input to the data acquisition unit 6a to an external server 101 described later via a communication module 7 described later. At the same time, the control unit 6 performs calculations using the input various data, pre-stored data, calculation formulas, etc. Then, the control unit 6 outputs the calculation result as a control command signal and is configured to control the operations of each part of the vehicle Ve respectively. In FIG. 1, an example in which one control unit 6 is provided for one vehicle Ve is shown, but a plurality of control units 6 may be provided for each device or equipment to be controlled, or for each control content.

[0035] The communication module 7 performs data transmission and reception between the control unit 6 of the vehicle Ve and a server 101 provided outside the vehicle Ve described later. The communication module 7 mounts, for example, a dedicated wireless communication system (not shown) called DCM [Data Communication Module] on the vehicle Ve, and transmits and receives various data between the control unit 6 and the server 101 using a dedicated communication line. In the embodiment of the present invention, general communication equipment (not shown) may be used to transmit and receive data using a general mobile communication line. Alternatively, for example, data transmission and reception may be performed using wired communication equipment installed in a vehicle Ve dealership, repair shop, etc.

[0036] In the time-series data processing apparatus according to the embodiment of the present invention, as described later, the control unit 6 transmits and receives data to and from a server 101 provided outside the vehicle Ve, and executes machine learning in cooperation with the server 101. For example, the deterioration state of the battery 4 as described above is estimated by machine learning using a neural network. At the same time, compression processing for reducing the amount of data and restoration processing for restoring the compressed data to a form supplied to machine learning are executed on the data used for the machine learning. Therefore, as shown in FIG. 2, the time-series data processing apparatus according to the embodiment of the present invention includes the in-vehicle control unit 6 and the server 101 installed outside the vehicle Ve.

[0037] Specifically, the control unit (ECU) 6 of the vehicle Ve has the data acquisition unit 6a and the transmission data creation unit 6b described above.

[0038] The data acquisition unit 6a acquires predetermined data necessary for generating learning data for machine learning for each vehicle Ve. The various data detected by the detection unit 5 described above are acquired as vehicle information for generating learning data for machine learning, that is, the original data of the time-series data in the embodiment of the present invention, as necessary. In the embodiment of the present invention, the time-series data is data that changes continuously over time. The original data of the time-series data is, for example, the data (analog data) as detected by the detection unit 5 as shown in FIG. 3(a). FIG. 3(a) shows data (voltage, internal resistance, temperature) related to the battery 4 of the vehicle Ve.

[0039] The transmission data creation unit 6b processes the various data acquired by the data acquisition unit 6a into data for transmission adapted to the communication module 7 as the original data of the time-series data used in machine learning, and transmits it to the server 101 via the communication module 7.

[0040] In addition, FIG. 2 shows a situation where two control units 6 transmit and receive data to and from the server 101, respectively. That is, it shows the control units 6 mounted on two vehicles Ve, respectively, and one server 101 set externally. The time-series data processing device in the embodiment of this invention executes machine learning using various data collected from the vehicle Ve. In order to improve the learning accuracy of the machine learning, it is desirable to collect as much data as possible acquired over as wide a range as possible. Therefore, the time-series data processing device in the embodiment of this invention is not limited to two vehicles Ve as shown in FIG. 2, and a large number of data are collected from the control units 6 mounted on a large number of vehicles Ve, respectively.

[0041] On the other hand, the server 101 installed outside the vehicle Ve has, for example, a data storage unit 101a, a data processing unit 101b, an arithmetic unit 101c, and a learning unit 101d.

[0042] The data storage unit 101a corresponds to the "storage unit" in the embodiment of this invention, and stores various data and information received from the control unit 6 of each vehicle Ve, and various data and information arithmetic-processed by the server 101, etc. in a storage device (not shown) or a storage medium (not shown) as a database related to vehicle information. In addition, it stores time-series data (original data) whose data amount has been reduced by compression processing performed by the data processing unit 101b described later in a storage device or a storage medium.

[0043] The data processing unit 101b corresponds to the "compression processing unit" and the "restoration processing unit" in the embodiment of this invention, and executes compression processing for reducing the data amount of a large number of time-series data that change continuously over time. In addition, it executes restoration processing on the data whose data amount has been reduced by compression processing.

[0044] As the compression process executed by the "compression processing unit" in this data processing unit 101b, as shown in FIG. 3(b), first, the original data of the time-series data stored in the data storage unit 101a is quantized (or discretized), and the first statistical index regarding the original data is calculated. For example, at least the average, standard deviation, and normal distribution function of the original data are calculated.

[0045] Also, data corresponding to the exclusion condition is excluded (downsampled) as missing values from the original data. The exclusion condition in this case corresponds to the "first condition" in the embodiment of this invention, and is, for example, a threshold value for discriminating data with low importance for a predetermined arithmetic process like the above-described machine learning. For example, it is a lower threshold value set as a lower limit value and an upper threshold value set as an upper limit value for the original data. Data below the lower threshold value and above the upper threshold value are data corresponding to the exclusion condition (first condition), and are excluded as missing values.

[0046] Also, the original data with the above-described missing values excluded is subjected to a predetermined data compression process to generate compressed data with a reduced data volume. In this case, existing data compression methods such as encryption and binarization are applied as the predetermined data compression process.

[0047] Then, as shown in FIG. 3(c), the above-described first statistical index and the compressed data are stored (memorized) in the "storage unit", that is, the data storage unit 101a, as storage data.

[0048] On the other hand, as the restoration process executed by the "restoration processing unit" in this data processing unit 101b, the stored data processed by the "compression processing unit" and stored in the data storage unit 101a is restored to data in a form that can be used by a predetermined arithmetic process, such as the above-described machine learning (learning data). Specifically, first, the stored data is read from the data storage unit 101a, and restored data obtained by subjecting the stored data to a predetermined process for restoring the stored data is generated. The predetermined process in this case is a data restoration process corresponding to the above-described data compression process, and existing data restoration methods such as decryption for encryption and so-called decompression for data compression are applied.

[0049] Also, as shown in (d) of FIG. 3, a second statistical index regarding the restored data is calculated. At the same time, from among the missing portions of the restored data including the traces from which the missing values have been excluded, positions corresponding to the data that satisfy a predetermined interpolation condition are specified as interpolation target portions. The interpolation condition in this case corresponds to the "second condition" in the embodiment of the present invention. For example, the data before and after the missing value in the time series direction of the original data is within the vicinity of the original data, that is, within a predetermined range including the lower threshold value of the original data or within a predetermined range including the upper threshold value of the original data. The positions of the missing portions corresponding to such vicinity of the original data are specified as the interpolation target portions.

[0050] Then, based on the first statistical index regarding the original data and the second statistical index regarding the restored data, an interpolation value to be applied to the interpolation target portion is calculated, and at the same time, the interpolation target portion is filled (interpolated) with the interpolation value to generate learning data in which the restored data approximates the original data.

[0051] Note that, as shown in FIG. 4, the time-series data processing apparatus in the embodiment of the present invention may include a data processing unit 6c in the control unit 6 of the vehicle Ve. In that case, the compression process of calculating the first statistical index and generating the compressed data as described above may be executed by the data processing unit 6c of the in-vehicle control unit 6. And the data processing unit 101b of the server 101 may be configured to execute only the above-described restoration process. That is, the data processing unit 6c of the control unit 6 may function as the "compression processing unit" in the embodiment of the present invention, and the data processing unit 101b of the server 101 may function as the "restoration processing unit" in the embodiment of the present invention.

[0052] The arithmetic unit 101c executes a predetermined arithmetic process based on the time-series data (learning data) that has been compressed by the data processing unit 101b as described above and then restored. For example, a well-known predetermined arithmetic process such as machine learning using a neural network is executed.

[0053] The learning unit 101d executes learning based on the result of the predetermined arithmetic process executed by the arithmetic unit 101c described above. For example, based on the result of machine learning as described above, the deterioration state of the battery 4 of the vehicle Ve, that is, the state of change over time of the performance of the battery 4, is estimated.

[0054] As described above, the time-series data processing method and apparatus in the embodiment of the present invention mainly aim to appropriately reduce and store the data volume while considering the necessity and importance of each individual data in a large amount of time-series data, and also to use it. For this purpose, the time-series data processing method and apparatus in the embodiment of the present invention are configured to execute the control shown in the following flowcharts.

[0055] The flowchart of FIG. 5 shows a compression process for performing a predetermined compression process on the original data of time-series data to reduce the data amount and convert it into storage data in a form that can be stored in a storage unit (data storage unit 101a), and a compression process step for executing the compression process. In the flowchart shown in FIG. 5, first, in step S101, in the control unit 6 of the vehicle Ve, various vehicle information including numerical data related to the battery 4 is acquired for each vehicle Ve.

[0056] Next, in step S102, the acquired various vehicle information is transmitted to the server 101 by wireless communication or the like. The vehicle information is acquired, for example, each time the ignition switch (not shown) is turned ON in each vehicle Ve. As described above, the transmission of vehicle information from the control unit 6 of the vehicle Ve to the external server 101 is not limited to wireless communication, and may be performed using a wired communication device or communication facility.

[0057] Subsequently, in step S103, the various vehicle information transmitted from each vehicle Ve and received is stored in the server 101. For example, it is grouped by a set (set data) for each vehicle (vehicle type) for a predetermined period. At this time, the numerical data related to the various vehicle information is quantized and processed into a form of time-series data that changes continuously over time. Then, it is temporarily stored in the server 101 as the original data of the time-series data.

[0058] In step S104, the first statistical index (or basic statistic) in the set data is calculated from the set data of the original data temporarily stored in the server 101 as described above. Specifically, at least the average value, standard deviation, and normal distribution function of the set data (original data) are calculated as the first statistical index. Also, the calculated first statistical index is stored in the server 101 together with the information (data count information) of each data in the set data (original data). Note that the timing of storing the data related to the first statistical index in the server 101 may be stored in the server 101 as storage data together with the compressed data in step S107 described later.

[0059] In step S105, a range is specified for the above set data (original data), and values outside that range are deleted. Specifically, data corresponding to the exclusion condition (the first condition) is excluded (thinned out) as missing values from the set data (original data). As described above, the exclusion condition in this case, that is, the first condition in the embodiment of the present invention, is a threshold value set to discriminate data with low importance for a predetermined arithmetic process (machine learning). And the data corresponding to the exclusion condition is data in the set data (original data) that is less than or equal to the lower threshold value set as the lower limit value and greater than or equal to the upper threshold value set as the upper limit value.

[0060] In step S106, predetermined (existing) data compression processing such as encryption or binarization is performed on the set data (original data) from which the above missing values have been excluded, and compressed data with a reduced data volume is generated.

[0061] Then, in step S107, the compressed data generated in step S106 above is stored in server 101 as, for example, data for battery degradation estimation learning, that is, data for storage for machine learning. As described above, in this step S107, the compressed data and the first statistical index may be stored in server 101 together as data for storage. When this step S107 is executed, then, the routine shown in the flowchart of FIG. 5 is terminated once.

[0062] The control content in each step shown in the flowchart of FIG. 5 corresponds to the compression processing step of the time series data processing method in the embodiment of the present invention. Specifically, the above-mentioned step S104 corresponds to the step of "calculating the first statistical index related to the original data" in the compression processing step, and step S105 corresponds to the step of "excluding, as missing values, the data that meet the predetermined first condition (exclusion condition) for discriminating the data with low importance for the predetermined arithmetic processing from the original data" in the compression processing step. Step S106 corresponds to the step of "performing compression processing on the original data with the missing values excluded to generate compressed data with the data volume reduced" in the compression processing step, and step S107 corresponds to the step of "storing the compressed data and the first statistical index in the storage unit as storage data" in the compression processing step.

[0063] As described above, in the time series data processing method and processing apparatus in the embodiment of the present invention, when compressing the original data of the time series data by the compression processing by the time series data processing apparatus and the compression processing step in the time series data processing method, for example, the first statistical indexes related to the original data, such as the average, standard deviation, and normal distribution function, are calculated. At the same time, the importance of each data is discriminated by the first condition (exclusion condition), and the data with low importance is excluded as a missing value. After reducing the missing values in this way, a predetermined compression processing is performed. Then, the compressed data after the compression processing is stored in the storage unit as storage data together with the above-mentioned first statistical index related to the original data. Therefore, it is possible to efficiently and effectively reduce and compress the data volume of the original data of the time series data while leaving the data with high importance, that is, while leaving the characteristics of the original data, and store it in the data storage unit 101a (storage unit).

[0064] On the one hand, the flowchart of FIG. 6 shows a restoration process of converting stored data into learning data in a form that can be used in a predetermined arithmetic process (machine learning), and a restoration process step for executing the restoration process. In the flowchart shown in FIG. 6, first, in step S201, for example, data for battery degradation estimation learning, that is, a set of stored data for machine learning, is prepared.

[0065] Next, in step S202, the first statistical indicators (such as mean, standard deviation, etc.) regarding the original data of the set of data, and the compressed data corresponding to the number of original data are read out.

[0066] In step S203, data restoration processing is performed on the read compressed data. For example, decryption of encrypted data or so-called decompression processing of compressed data is executed to generate restored data.

[0067] In step S204, the second statistical indicators regarding the restored data are calculated from the restored (decrypted) data. For example, the mean, standard deviation of the restored data, and the normal distribution function (discrete data with specified number) having these mean and standard deviation are calculated.

[0068] In step S205, based on the calculated normal distribution function, data corresponding to missing values (values less than or equal to "mean - standard deviation") are extracted. Specifically, among the missing locations of the original data, the position corresponding to the data that meets the interpolation condition (second condition) is specified as the interpolation target location, and the interpolation value to be applied to the interpolation target location is calculated based on the second statistical indicators regarding the restored data. In this case, the interpolation value is extracted from the random numbers obtained based on the normal distribution function of the restored data among the calculated second statistical indicators.

[0069] In step S206, the interpolation value extracted as described above is applied to the missing location. That is, the interpolation target location of the missing location is interpolated by the interpolation value. That is, the interpolation target location (missing value) is approximately supplemented by the interpolation value.

[0070] Then, in step S207, the interpolation target portion is interpolated with the interpolation value as described above, so that learning data approximating the restored data to the original data is generated. That is, the learning data is generated so that the shape of the normal distribution in the first statistical index regarding the original data is approximated to the shape of the normal distribution in the second statistical index regarding the restored data. When this step S207 is executed, then, the routine shown in the flowchart of FIG. 6 is temporarily terminated.

[0071] The control contents in each step shown in the flowchart of FIG. 6 correspond to the restoration process step of the time series data processing method in the embodiment of the present invention. Specifically, the above-described step S202 corresponds to the step of "reading the storage data (compressed data and the first statistical index) from the storage unit" in the restoration process step, step S203 corresponds to the step of "generating restored data obtained by subjecting the storage data to a predetermined process for restoring the storage data" in the restoration process step, step S204 corresponds to the step of "calculating the second statistical index regarding the restored data" in the restoration process step, step S205 corresponds to the step of "identifying, as the interpolation target portion, the position corresponding to the data that satisfies a predetermined second condition (interpolation condition) from among the missing portions of the restored data including the trace from which the missing values are excluded" in the restoration process step, and steps S206 and S207 correspond to the step of "calculating an interpolation value to be applied to the interpolation target portion based on the first statistical index and the second statistical index, and supplementing the interpolation target portion of the restored data with the calculated interpolation value to generate learning data approximating the restored data to the original data" in the restoration process step.

[0072] As described above, in the time-series data processing method and processing apparatus according to the embodiment of the present invention, in the restoration processing by the time-series data processing apparatus and the restoration processing step in the time-series data processing method, the first statistical index obtained from the original data and the second statistical index obtained from the restored data are used, and learning data approximated to the original data is generated. Specifically, when restoring the stored data (compressed data) stored in a compressed manner, the missing values in the stored data are interpolated so that the shape of the normal distribution of the original data and the shape of the normal distribution of the learning data to be restored are approximated. Therefore, it is possible to restore learning data that appropriately reproduces the characteristics of the compressed original data.

[0073] As described above, according to the time-series data processing method and processing apparatus according to the embodiment of the present invention, considering the necessity and importance of each individual data in a large amount of time-series data, it is possible to efficiently reduce the data amount while leaving important information. Therefore, it is possible to appropriately store a large amount of time-series data without straining the storage capacity of the data storage unit 101a (storage unit). And the time-series data stored in a compressed manner can be restored to be well approximated to the original data, and can be appropriately used, for example, as learning data for machine learning.

Explanation of Reference Numerals

[0074] 1 Driving force source (POWER) 2 Driving wheels 3 Starter motor 4 Battery (power storage device) 5 Detection unit 5a Battery voltage sensor 5b Travel distance sensor 5c Outside air temperature sensor 5d Battery temperature sensor 5e Battery resistance sensor 5f Counter 5g Vehicle speed sensor (or wheel speed sensor) 5h Rotation speed sensor 6 Control unit (ECU) Data acquisition unit (of the control unit) Transmission data creation unit (of the control unit) Data processing unit (compression processing unit) (of the control unit) 7 Communication module (DCM) 8 Engine (driving power source) 9 Transmission 10 Differential gear 11 Drive shaft 101 Server 101a Data storage unit (memory unit) (of the server) 101b Data processing unit (compression processing unit, decompression processing unit) (of the server) 101c Arithmetic unit (of the server) 101d Learning unit (of the server) Ve Vehicle

Claims

1. A method for processing time-series data for reducing the data volume of a large number of time-series data that change continuously over time, a compression processing step of performing a predetermined compression process on the original data of the quantized time-series data, reducing the data volume, and obtaining storage data in a form that can be stored in a predetermined storage unit; a restoration processing step of performing a restoration process on the storage data to obtain learning data in a form that can be used in a predetermined arithmetic process; comprising: the compression processing step includes: a step of calculating a first statistical index related to the original data; a step of excluding, as missing values, data that meet a predetermined first condition from the original data; a step of performing the compression process on the original data from which the missing values have been excluded, and generating compressed data with reduced data volume; a step of storing the compressed data and the first statistical index in the storage unit as the storage data; and has: the restoration processing step includes: a step of reading the storage data from the storage unit; a step of generating restored data by performing a predetermined process for restoring the storage data on the storage data; a step of calculating a second statistical index related to the restored data; a step of specifying, as interpolation target locations, positions corresponding to data that meet a predetermined second condition among the missing locations of the restored data including traces of the excluded missing values; a step of calculating an interpolation value to be applied to the interpolation target locations based on the first statistical index and the second statistical index, filling the interpolation target locations with the interpolation values, and generating the learning data in which the restored data approximates the original data; and has: A method for processing time-series data, characterized by the above.

2. The method for processing time-series data according to claim 1, wherein The data that meets the first condition is the data in the original data that is less than or equal to the lower threshold value set as the lower limit value and greater than or equal to the upper threshold value set as the upper limit value. The data that meets the second condition is the data in the original data where the data before and after the missing value in the time series direction is within a predetermined range including the lower threshold value or within a predetermined range including the upper threshold value. The interpolated value is extracted from a random number obtained based on the normal distribution function of the restored data. A method for processing time series data, characterized by the above.

3. A method for processing time series data according to claim 2, The arithmetic processing is machine learning that performs learning based on a large number of input data and makes predictions or judgments based on the results of the learning. The learning data is the input data in the machine learning. A method for processing time series data, characterized by the above.

4. A method for processing time series data according to claim 3, The original data is data related to a power storage device mounted on a vehicle. The machine learning predicts the change over time of the power storage device. A method for processing time series data, characterized by the above.

5. A time series data processing device that reduces the data volume of a large number of time series data that changes continuously over time, Among the original data of the quantized time series data, the original data from which the data that meets a predetermined exclusion condition is excluded as a missing value is subjected to a predetermined compression process to reduce the data volume, and a restoration process is performed on the stored data stored in a predetermined storage device or storage medium to obtain learning data in a form that can be used in a predetermined arithmetic process. It is equipped with a restoration processing unit. The restoration processing unit, Reads the stored data from the storage device or the storage medium, Generates restored data obtained by performing a predetermined process for restoring the stored data on the stored data. Calculate the second statistical index related to the restored data, From among the missing locations of the restored data including the trace where the missing values are excluded, identify the positions corresponding to the data that meet a predetermined interpolation condition as the interpolation target locations, Based on the first statistical index related to the original data and the second statistical index, calculate an interpolation value to be applied to the interpolation target locations, and supplement the interpolation target locations with the interpolation values to generate the learning data in which the restored data approximates the original data A time-series data processing device characterized by the above.

6. A time-series data processing device according to claim 5, The data that meets the exclusion condition is data in the original data that is less than or equal to a lower threshold value set as a lower limit value and greater than or equal to an upper threshold value set as an upper limit value, The data that meets the interpolation condition is data within a predetermined range including the lower threshold value or within a predetermined range including the upper threshold value for the data before and after the missing value in the time series direction of the original data, The interpolation value is extracted from a random number obtained based on the normal distribution function of the restored data A time-series data processing device characterized by the above.

7. A time-series data processing device according to claim 6, The arithmetic processing is machine learning that performs learning based on a large number of input data and makes a prediction or judgment based on the result of the learning, The learning data is the input data in the machine learning, The original data is data related to a power storage device mounted on a vehicle, The machine learning predicts the change over time of the power storage device A time-series data processing device characterized by the above.

Citation Information

Patent Citations

  • Arithmetic unit for neural network

    JP1993265997A

  • Compression method of time series data and compression device

    JP2012010319A

  • Data collection system and method for controlling data collection system

    JP2018142158A

  • Complement device, complement method and its program of time series data of instant heart rate

    JP2018201787A

  • Tool replacement timing management system

    JP2020163493A