Information Processing Apparatus, Information Processing Method, and Information Processing Program

The information processing apparatus addresses the challenge of grouping multivariate time series data by training an RNN-LSTM inference device with divergence calculations and using Attention Weight for effective two-group classification, enabling visualization of contributing time and variables.

JP7694485B2Active Publication Date: 2025-06-18TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022109976
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2025-06-18
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

Existing technologies face challenges in effectively grouping multivariate time series data due to the black-box nature of machine learning systems and the difficulty in handling time series data using Convolutional Neural Networks (CNNs).

Method used

An information processing apparatus that includes a first divergence calculation unit to calculate the divergence of second group data from first group data, a learning unit that trains an RNN-LSTM inference device using this divergence, a second divergence calculation unit to calculate the divergence of new analysis target data from the first group data, and a classification unit that performs two-group classification using the trained inference device.

Benefits of technology

Enables effective two-group classification of multivariate time series data by using RNN-LSTM with Attention Weight, allowing for visualization of time and variables with significant contribution to the classification, thereby overcoming the limitations of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694485000001
    Figure 0007694485000001
  • Figure 0007694485000002
    Figure 0007694485000002
  • Figure 0007694485000003
    Figure 0007694485000003
Patent Text Reader

Abstract

To provide an information processing apparatus, an information processing method, and an information processing program which classify multivariate time series data into groups.SOLUTION: A degree of deviation of a B group sample set different from an A group sample set, from the A group sample set is calculated, and the degree of deviation of the B group sample set from the A group sample set is used to train an RNN-LSTM (Recurrent Neural Network-Long Short Term Memory) being a reasoner. After the training, a degree of deviation of a new analysis object sample from the A group sample set is calculated, and the degree of deviation of the new analysis object sample from the A group sample set is applied to the trained reasoner to perform two-group classification for determining which of the A group sample set and the B group sample set the new analysis object sample belongs to.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program for grouping multivariate time series data.

Background Art

[0002] Group classification of multivariate time series data using a machine learning system has been performed, but the machine learning system has a problem that the process leading to the recognition result is black-boxed.

[0003] Patent Document 1 discloses an invention of an inference apparatus that can infer to which of a predetermined class an input image belongs and output, together with the inference result, a feature amount that is the basis of the inference.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, since the invention described in Patent Document 1 uses an input data as an image and a Convolutional neural network (CNN) as a neural network, it has been difficult to handle time series data.

[0006] In consideration of the above facts, an object of the present invention is to provide an information processing apparatus, an information processing method, and an information processing program for grouping multivariate time series data.

Means for Solving the Problems

[0007] To achieve the above object, the information processing apparatus according to claim 1 includes: a first divergence calculation unit that calculates a divergence of second group data from the first group data, which is different from the first group data; a learning unit that trains an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory), which is an inference device, using the divergence calculated by the first divergence calculation unit; a second divergence calculation unit that calculates a divergence of new analysis target data from the first group data; and a classification unit that applies the divergence calculated by the second divergence calculation unit to the trained inference device to perform two-group classification of which of the first group data and the second group data the new analysis target belongs to.

[0008] In the information processing apparatus according to claim 1, the input to the inference device is the divergence from either of the two groups. Since an increase in the input variable to the inference device indicates an increase in the divergence from one group, two-group classification of the data can be performed based on such a matter.

[0009] Further, according to the information processing apparatus according to claim 1, by adopting an RNN-LSTM that outputs an Attention Weight indicating the importance of the time when inferring the data input after group classification, two-group classification of multivariate time series data becomes possible.

[0010] In the information processing apparatus according to claim 2, the first divergence calculation unit calculates the divergence of the second group data from the first group data based on a Gaussian mixture model of the first group data, and the second divergence calculation unit calculates the divergence of the new analysis target data from the first group data based on the Gaussian mixture model of the first group data, as in the information processing apparatus according to claim 1.

[0011] According to the information processing apparatus according to claim 2, the divergence between the two groups of data is calculated by a Gaussian mixture model, which is a probability distribution obtained by overlapping normal distribution curves.

[0012] The information processing apparatus according to claim 3, wherein the learned inference device outputs an estimated probability indicating the probability that the new analysis target belongs to either the first group of data or the second group of data, and the classification unit is based on the divergence calculated by the second divergence calculation unit and the estimated probability. It estimates whether the new analysis target belongs to either the first group of data or the second group of data, and after the estimation, based on the weighting calculated using the learned inference device and the divergence calculated by the second divergence calculation unit, it calculates the contribution degree of the new analysis target to the classification into the second group of data.

[0013] According to the information processing apparatus described in claim 3, at the input time that was important for the two-group classification indicated by the weighting, the contribution degree of the input variables with a large divergence in the two-group classification is calculated. As a result, it becomes possible to visualize the time and variables with a large contribution degree in the two-group classification.

[0014] To achieve the above object, the information processing method according to claim 4 includes a first divergence calculation step of calculating the divergence of the second group of data from the first group of data, which is different from the first group of data, and using the divergence calculated in the first divergence calculation step. A learning step of training an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory), which is an inference device, a second divergence calculation step of calculating the divergence of the new analysis target from the first group of data, and applying the divergence calculated in the second divergence calculation step to a learned inference device to perform a two-group classification of whether the new analysis target belongs to the first group of data or the second group of data.

[0015] In the information processing method according to claim 4, the input to the inference device is the divergence from either of the two groups. Since the increase in the input variable to the inference device represents the increase in the divergence from one of the groups, the two-group classification of the data can be performed based on such matters.

[0016] In addition, according to the information processing method described in claim 4, by adopting an RNN-LSTM that outputs an Attention Weight indicating the importance of the time when inferring the data input after group classification, group classification of multivariate time-series data becomes possible.

[0017] To achieve the above object, the information processing program described in claim 5 causes a computer to function as a first divergence calculation unit that calculates the degree of divergence of second group data different from the first group data from the first group data, a learning unit that uses the degree of divergence calculated by the first divergence calculation unit to train an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory) as an inference device, a second divergence calculation unit that calculates the degree of divergence of new analysis target data from the first group data, and a classification unit that applies the degree of divergence calculated by the second divergence calculation unit to the trained inference device to perform two-group classification of whether the new analysis target belongs to either the first group data or the second group data.

[0018] The information processing program described in claim 5 uses the degree of divergence from either of the two groups as the input to the inference device. Since an increase in the input variable to the inference device indicates an increase in the divergence from one group, two-group classification of the data can be performed based on such matters.

[0019] In addition, according to the information processing program described in claim 5, by adopting an RNN-LSTM that outputs an Attention Weight indicating the importance of the time when inferring the data input after group classification, group classification of multivariate time-series data becomes possible.

Advantages of the Invention

[0020] As described above, according to the information processing apparatus, information processing method, and information processing program according to the present invention, group classification of multivariate time-series data becomes possible.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Mode for Carrying Out the Invention

[0022] Hereinafter, the information processing apparatus 10 according to the present embodiment will be described with reference to FIG. 1. As shown in FIG. 1, the information processing apparatus 10 includes a computer 30. The computer 30 includes a CPU 32, a ROM 34, a RAM 36, and an input / output port 38. As an example, the computer 30 is preferably a model that can execute advanced arithmetic processing at high speed, such as an engineering workstation or a supercomputer.

[0023] In the computer 30, the CPU 32, the ROM 34, the RAM 36, and the input / output port 38 are connected to each other via various buses such as an address bus, a data bus, and a control bus. To the input / output port 38, as various input / output devices, a display 40, a mouse 42, a keyboard 44, a hard disk (HDD) 46, and a disk drive 50 for reading information from various disks (for example, CD-ROM, DVD, etc.) 48 are respectively connected.

[0024] In addition, a network 52 is connected to the input / output port 38, and information can be exchanged with various devices connected to the network 52. In the present embodiment, a data server 56 to which a database (DB) 54 is connected is connected to the network 52, and information can be exchanged with the DB 54. The DB 54 stores multivariate time-series data to be evaluated and the like. The multivariate time-series data to be evaluated is stored in the DB 54 via the network 52 from the input device 60. If the input device 60 is connected to the network 52, the computer 30 may acquire it from the input device 60.

[0025] In the present embodiment, although it is described that the DB 54 connected to the data server 56 stores multivariate time-series data to be evaluated and the like, the information of the DB 54 may be stored in an external storage device such as the HDD 46 built in the computer 30 or an external hard disk.

[0026] A program related to machine learning is installed in the HDD 46 of the computer 30. In the present embodiment, the CPU 32 executes the program, whereby machine learning is started and a learned model based on the machine learning is constructed. The CPU 32 displays the processing result by the program on the display 40.

[0027] There are several methods for installing the program related to the machine learning of the present embodiment in the computer 30. For example, the program is stored together with a setup program in a CD-ROM, a DVD, or the like, the disk is set in the disk drive 50, and the setup program is executed for the CPU 32 to install the program in the HDD 46. Alternatively, the program may be installed in the HDD 46 by communicating with another information processing device connected to the computer 30 via a public telephone line or the network 52.

[0028] By executing a program related to machine learning, the CPU 32 has a divergence degree calculation function for calculating the divergence degree indicating the degree of divergence of data in two groups of multivariate time series data, a learning function for performing machine learning by providing the calculated divergence degree to a mathematical model as an inference device, and a specified probability output function for outputting a specified probability indicating the probability that input data belongs to either one of the above two multivariate time series data. By executing the program related to machine learning having these functions, the CPU 32 functions as a divergence degree calculation unit, a learning unit, and a specified probability output unit.

[0029] After learning, the CPU 32 has a divergence degree calculation function for calculating the divergence degree of the input new multivariate time series data from the multivariate time series data as reference data, a classification function for performing two-group classification of the input data using the learned inference device, and a contribution degree calculation function for calculating the contribution degree of each variable to the classification. By executing the program having these functions, the CPU 32 functions as a divergence degree calculation unit, a classification unit, and a contribution degree calculation unit.

[0030] FIG. 2 is a block diagram showing an example of the processing of the information processing apparatus 10 according to the present embodiment. As shown in FIG. 2, the processing of the information processing apparatus 10 according to the present embodiment includes a learning-time block 110 for performing machine learning and an inference-time block 130 using the learned model. Further, the processing of the information processing apparatus 10 according to the present embodiment includes a divergence degree calculation block 120 for calculating the divergence degree indicating the degree of divergence of data in two groups of multivariate time series data and a contribution degree calculation block 140 for calculating the contribution degree of each variable to the two-group classification.

[0031] In step S100 in the learning block 110, a set of A-group samples of multivariate time-series data serving as reference data (hereinafter abbreviated as "A-group") is input. As an example, the A-group is multivariate time-series data related to the straight-ahead start behavior of a vehicle by a group of drivers with normal cognitive ability. Specifically, the A-group is 8-dimensional multivariate time-series data with variables of vehicle speed, longitudinal acceleration, longitudinal jerk, lateral acceleration, lateral jerk, yaw rate, yaw angular acceleration, and yaw angular jerk in the straight-ahead start behavior of the vehicle, and is time-series data for 6.4 seconds (64 time points) in 0.1-second units after starting from the stopped state of the vehicle.

[0032] In step S102, as A-group distribution learning, for example, the changes over time for each variable are graphed. Then, in step S104, as the A-group distribution shape, for example, for each variable and each of the 64 time points, a normal distribution (Gaussian distribution) is calculated with the horizontal axis being a random variable indicating one variable such as vehicle speed and the vertical axis being the probability density indicating the appearance frequency of a plurality of drivers belonging to the A-group accompanying the change of the random variable shown on the horizontal axis.

[0033] In step S106, the degree of deviation from the A-group is calculated by comparing the A-groups. When comparing the A-groups, the same data is compared, so theoretically, the degree of deviation indicates 0. In this embodiment, the degree of deviation is calculated by a Gaussian mixture model (GMM). Although the details of the GMM will be described later, specifically, it is a probability distribution obtained by overlapping a plurality of normal distribution curves as shown in FIG. 3.

[0034] In step S108, a set of B-group samples of multivariate time-series data (hereinafter abbreviated as "B-group") is input. As an example, the B-group is multivariate time-series data related to the straight-ahead start behavior of a vehicle by a group of drivers with reduced cognitive ability. Specifically, the B-group is, like the A-group, 8-dimensional multivariate time-series data with variables of vehicle speed, longitudinal acceleration, longitudinal jerk, lateral acceleration, lateral jerk, yaw rate, yaw angular acceleration, and yaw angular jerk in the straight-ahead start behavior of the vehicle, and is time-series data for 6.4 seconds (64 time points) in 0.1-second units after starting from the stopped state of the vehicle.

[0035] In step S110, the GMM calculates the degree of divergence of group B from group A. FIG. 3 is a schematic diagram showing an example of GMM. As shown in FIG. 3, by adjusting the mean (median) and covariance (the degree of spread of the peak) in each of the normal distribution 0 and the normal distribution 1, and adding them with appropriate weight coefficients (linear combination), the GMM curve as shown in FIG. 3 is obtained.

[0036] The horizontal axis in FIG. 3 is a probability variable indicating a variable 1 such as the vehicle speed at a specific time (time 1), and the vertical axis is the probability density (likelihood) indicating the appearance frequency of a plurality of drivers belonging to group A accompanying the change of the probability variable shown on the horizontal axis.

[0037] In the present embodiment, the likelihood in the GMM indicated by the distribution of the values of the same variable of each data belonging to group A at the same time as the value of the variable 1 of a certain piece of data 1 at a certain time 1 is defined as the degree of divergence from group A of the same variable of that data at the same time. Then, the calculated degrees of divergence are each expressed in logarithm, arranged in time order for each of the aforementioned 8 dimensions, and made into time series data of 8 dimensions and 64 time points. In the present embodiment, such time series data of 8 dimensions and 64 time points is referred to as divergence degree time series data.

[0038] In step S112, the divergence degree time series data of group A with respect to group A is prepared as described above. In step S114, the divergence degree time series data of group B with respect to group A is prepared as described above.

[0039] The information processing apparatus 10 according to the present embodiment classifies input data into either group A or group B, and uses an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory) having an Attention mechanism indicating the input time that was important for inference. Then, as preprocessing for inference, for each sample (such as group B), the degree of divergence from one group (group A) is obtained for each variable, and this is collectively input to the RNN-LSTM as divergence degree time series data. Using the divergence degree time series data as the input data of the RNN-LSTM is common during both learning and inference.

[0040] The fact that the input variable of the divergence time series data input to the RNN-LSTM increases indicates that the divergence from Group A becomes larger. Therefore, in the two-group classification, if it is determined that the sample does not belong to Group A based on the divergence, since the sample belongs to Group B, the magnitude of the input variable will contribute to the classification into Group B. After the two-group classification, the RNN-LSTM outputs a weight indicating the importance of the time when inferring the input data called Attention Weight. Therefore, based on the value of such Attention Weight, the input time that was important for the two-group classification can be known. Furthermore, it can be estimated that the input variable with a large divergence at the input time that was important for the two-group classification contributes to the classification into Group B.

[0041] In step S116, machine learning is performed using the calculated divergence time series data. As described above, in this embodiment, RNN-LSTM is adopted as the mathematical model which is the inference device. In step S116, as an example, the label [1,0] is associated with the starting behavior related to Group A, and the label [0,1] is associated with the starting behavior related to Group B, and the divergence time series data is input.

[0042] In step S118, by using the softmax function as the activation function, the estimated probability of whether the input in step S116 is from Group A or Group B is output, and the process in the learning block 110 is terminated.

[0043] Hereinafter, a series of processes in the inference block 130 will be described. In step S120, the analysis target sample is input.

[0044] In step S122, the divergence of the analysis target sample from Group A is calculated. The calculation of the divergence is performed using GMM as described above.

[0045] In step S124, similar to step S114, the divergence time series data of the analysis target sample with respect to Group A is prepared.

[0046] In step S126, binary classification is performed using the aforementioned estimated probability to determine whether the sample to be analyzed belongs to group A or group B. In step S128, the classification result is stored in the RAM 36 or the HDD 46.

[0047] In step S130, an Attention Weight, which is a weight indicating the importance of the time when inferring the two-group classification of the sample to be analyzed, is output. The Attention Weight is time-series data with the same 64-time series length as the deviation degree time-series data that is the input data.

[0048] In step S132, the Attention Weight is multiplied by each variable of the deviation degree time-series data. In step S134, the product obtained in step S132 is used as the degree suggesting that the input data belongs to group B, that is, the contribution degree in the two-group classification for each variable and each time.

[0049] FIG. 4 is a schematic diagram showing the contribution degree for each variable and each time in the starting behavior classified as group B. In FIG. 4, it shows that two variables, the yaw angular acceleration and the yaw angular jerk, particularly contribute to the classification of the sample to be analyzed into group B immediately after starting and when a certain amount of time (approximately 3.5 seconds) has passed after starting.

[0050] In step S136, the classification result stored in step S128 and the contribution degree obtained in step S134 are output to end the series of processes.

[0051] As described above, in the present embodiment, in the two-group classification, the input to the inferencer is the deviation degree from either of the two groups. Since an increase in the input variable to the inferencer represents an increase in the deviation from one of the groups, the two-group classification of the data is performed based on such matters.

[0052] Also, in the present embodiment, by adopting an RNN-LSTM, which is a type of neural network, as the inferencer, it becomes possible to classify groups of multivariate time-series data.

[0053] The RNN-LSTM outputs an Attention Weight that indicates the importance of time when inferring the data input after group classification. In the present embodiment, using the aforementioned divergence degree and Attention Weight, the contribution degree of the input variables with a large divergence degree in the two-group classification at the input time that was important for the two-group classification indicated by the Attention Weight is calculated. As a result, it becomes possible to visualize the time and variables with a large contribution degree in the two-group classification.

[0054] In the present embodiment, as an example, time-series data such as vehicle speed and acceleration in the straight-ahead start behavior of a vehicle was handled. Since information related to driving an automobile is time-series data, it can be utilized, for example, for state estimation such as a drive monitor. Further, the present embodiment is not limited to vehicle-related matters and can widely handle classification problems related to time-series data.

[0055] Note that in each of the above embodiments, the processing executed by the CPU by reading software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) whose circuit configuration can be changed after manufacturing, such as an FPGA (Field-Programmable Gate Array), and a dedicated electric circuit which is a processor having a circuit configuration dedicated to executing specific processing, such as an ASIC (Application Specific Integrated Circuit). Further, the processing may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a plurality of FPGAs, and a combination of a CPU and an FPGA, etc.). Further, the hardware structure of these various processors is, more specifically, an electric circuit combining circuit elements such as semiconductor elements.

[0056] In addition, in each of the above embodiments, although the mode in which the program is pre-stored (installed) in the disk drive 50 or the like has been described, the present invention is not limited to this. The program may be provided in a form stored in a non-transitory storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), and a USB (Universal Serial Bus) memory. Further, the program may be in a form downloaded from an external device via a network.

[0057] Note that the "first group of data" described in the claims corresponds to the "set of A-group samples" described in the detailed description of the invention, the "second group of data" described in the claims corresponds to the "set of B-group samples" described in the detailed description of the invention, the "first divergence calculation unit" described in the claims corresponds to the "divergence calculation unit (during learning)" described in the detailed description of the invention, the "learning unit" described in the claims corresponds to the "learning unit" described in the detailed description of the invention, the "second divergence calculation unit" described in the claims corresponds to the "divergence calculation unit (after learning)" described in the detailed description of the invention, and the "classification unit" described in the claims corresponds to the "classification unit" described in the detailed description of the invention, respectively.

[0058] (Additional item 1) A memory, At least one processor connected to the memory, Including, The processor is Calculate the divergence of a second group of data different from the first group of data from the first group of data, Use the divergence of the second group of data from the first group of data to train an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory) which is an inference device, Calculate the divergence of a new analysis target from the first group of data, Applying the degree of deviation from the first group of data of the new analysis target to a learned inference device to perform two-group classification of whether the new analysis target belongs to either the first group of data or the second group of data. An information processing apparatus configured as follows.

Explanation of symbols

[0059] 10 Information processing apparatus 30 Computer 32 CPU 34 ROM 36 RAM 38 Input / output port 40 Display 42 Mouse 44 Keyboard 50 Disk drive 52 Network 56 Data server 60 Input device 110 Learning block 120 Deviation calculation block 130 Inference block 140 Contribution calculation block

Claims

1. A first divergence calculation unit that calculates the degree of divergence of second group data different from the first group data from the first group data; A learning unit that uses the degree of divergence calculated by the first divergence calculation unit to train an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory) which is an inference device; A second divergence calculation unit that calculates the degree of divergence of new analysis target data from the first group data; A classification unit that applies the degree of divergence calculated by the second divergence calculation unit to a trained inference device to perform two-group classification of whether the new analysis target belongs to either the first group data or the second group data; An information processing apparatus including the above.

2. The first divergence calculation unit calculates the degree of divergence of the second group data from the first group data based on a Gaussian mixture model of the first group data, The second divergence calculation unit calculates the degree of divergence of the new analysis target data from the first group data based on a Gaussian mixture model of the first group data. The information processing apparatus according to claim 1.

3. The trained inference device outputs an estimated probability indicating the probability that the new analysis target belongs to either the first group data or the second group data, The classification unit estimates whether the new analysis target belongs to either the first group data or the second group data based on the degree of divergence calculated by the second divergence calculation unit and the estimated probability, and after the estimation, calculates the contribution degree to the classification of the new analysis target into the second group data based on the weighting calculated using the trained inference device and the degree of divergence calculated by the second divergence calculation unit. The information processing apparatus according to claim 1 or 2.

4. A first divergence calculation step of calculating the degree of divergence of second group data different from the first group data from the first group data; A learning step of training an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory), which is an inference device, using the degree of deviation calculated in the first degree-of-deviation calculation step; A second degree-of-deviation calculation step of calculating the degree of deviation from the first group of data of a new analysis target; A classification step of applying the degree of deviation calculated in the second degree-of-deviation calculation step to a trained inference device to perform two-group classification as to which of the first group of data and the second group of data the new analysis target belongs to; An information processing method including the above.

5. A computer Function as a first degree-of-deviation calculation unit that calculates the degree of deviation from the first group of data of a second group of data different from the first group of data, a learning unit that trains an RNN-LSTM (Recurrent Neural Network - Long Short Term Memory), which is an inference device, using the degree of deviation calculated by the first degree-of-deviation calculation unit, a second degree-of-deviation calculation unit that calculates the degree of deviation from the first group of data of a new analysis target, and a classification unit that applies the degree of deviation calculated by the second degree-of-deviation calculation unit to a trained inference device to perform two-group classification as to which of the first group of data and the second group of data the new analysis target belongs to. An information processing program.

Citation Information

Patent Citations

  • Inference device, inference method and program

    JP2019082883A

  • Control state monitoring system and program

    JP2021022290A

  • Abnormality detection system, abnormality detection method, and abnormality detection program

    JP2021056927A