Change Point Score Calculation Device, Change Point Score Calculation Method, and Program
By calculating centroid distance matrices and weighting squared errors, the method addresses the issue of inaccurate change point scores due to divergent patterns, resulting in more precise change point detection.
Patent Information
- Application Number
- JP2024517649
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2042-04-26
AI Technical Summary
Existing change point detection methods fail to consider the degree of separation between cluster transition patterns, leading to inaccurate change point scores when new patterns emerge that are far from or similar to past observations.
Introduce a mechanism that calculates the centroid distance matrix for cluster pairs and weights the squared error of residence probabilities to enhance change point scores when new patterns diverge significantly from past observations.
This approach promotes more precise change point detection by increasing the score when new patterns are significantly different from past or current observations, enhancing the accuracy of change point detection.
Smart Images

Figure 0007709660000003 
Figure 0007709660000004 
Figure 0007709660000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to calculating a change point score considering distance.
Background Art
[0002] A technique for detecting a change point in the system state of a system composed of one or more devices using time series data representing the system state at each point in time has been conventionally known. Here, the "system state" refers to the operating state of the system represented by quantitative variables such as "number of accesses" and "number of users".
[0003] As a technique for detecting a change point for time series data to which no correct label regarding the occurrence position of the change point is assigned, the technique described in Non-Patent Document 1 is known.
[0004] The method of Non-Patent Document 1 is a method that extends the technique of detecting a change point by clustering, and since it is a clustering-based method, the target time series is not subject to constraints such as stationarity constraints and independent and identically distributed constraints. Further, the method of Non-Patent Document 1 can be said to be a method that introduces the concept of the time axis in that after clustering the time windows at each point in time of the time series data, the clusters assigned to each point in time are tracked in the time axis direction to extract the transition pattern. Furthermore, the method of Non-Patent Document 1 sets a past period and a current period that are sufficiently longer than the time window, compares the distributions of the cluster transition patterns in both periods, and calculates a change point score. That is, it calculates a change point score for interval data having a certain time width rather than snapshot data for each point in time, and is a technique capable of detecting a change point including a change in the time series pattern.
[0005] The change point detection device proposed in Non-Patent Document 1 specifically has an input unit that inputs time-series data representing the system state at each point in time of a system composed of one or more devices, and the time-series data is composed of data in the dimension of the number of devices constituting the system × the number of items representing the state of the device, and a time window generation unit that generates conversion data by converting the time-series data at each point in time from data in the dimension of the number of devices × the number of items into data in the dimension of the number of devices × the number of items × the time window length, and a function (cluster part, cluster transition series creation part, cluster transition tensor calculation part, change point score calculation part, detection part) that detects a change point when the change point score of the system state calculated based on the conversion data at each point in time exceeds a preset threshold value.
[0006] Further, the change point score calculation unit executes by the following mean squared error as a calculation method of the change point score, that is, the distance between the cluster transition tensors in the past period and the current period. dist(D1,D2)=(Σ l=1 L Σ m=1 M (d2 c1,···,cL -d1 c1,···,cL ) 2 / M L ) 1 / 2 Here, D1 and D2 are the cluster transition tensors in the past period and the current period, respectively, and d1 c1,···,cL , d2 c1,···,cL are the elements of the tensors D1 and D2 in which the residence probabilities of the cluster transition patterns {c1, c2, ···, cL} are stored, respectively.
Prior Art Documents
Non-Patent Documents
[0007]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] However, in the above method for calculating the distance between tensors, for example, when there is a cluster transition pattern that has not been observed in the past period but is newly observed in the current period, even if the newly observed cluster transition pattern is far from or similar to other cluster transition patterns that have been observed in the past, the result of the distance calculation will be the same. Conversely, when there is a cluster transition pattern that was observed in the past period but is not observed in the current period, even if the cluster transition pattern observed in the past is far from or similar to other cluster transition patterns currently observed, the result of the distance calculation will still be the same. When a cluster transition pattern far from the past observations is newly observed, or when a cluster transition pattern far from the currently observed patterns was observed in the past, it is desirable for the change point score, that is, the distance between the cluster transition tensors in the past period and the current period, to increase. However, the above method for calculating the distance between tensors has a problem that it cannot consider the degree of separation (distance) between such cluster transition patterns.
[0009] The present invention has been made in view of the above points, and aims to introduce a mechanism that promotes an increase in the change point score when a cluster transition pattern far from the past observations is newly observed, or when a cluster transition pattern far from the currently observed patterns was observed in the past, and to achieve more refined change point detection.
Means for Solving the Problems
[0010] To achieve the above object, the invention according to claim 1 is a centroid coordinate input unit that inputs centroid coordinates that are the centers of all clusters assigned to data of the dimension of the number of devices × the number of items × the time window length at each time point constituting the time series data of the past period and the current period generated by the clustering unit, a cluster transition pattern input unit that inputs all cluster transition patterns that appeared in the past period and the current period extracted by the cluster transition tensor calculation unit, a cluster transition tensor input unit that inputs the cluster transition tensors of the past period and the current period calculated by the cluster transition tensor calculation unit, a centroid distance matrix calculation unit that calculates a centroid distance matrix for all cluster pairs based on the centroid coordinates of all clusters input by the centroid coordinate input unit, a cluster transition pattern distance matrix calculation unit that calculates a distance matrix for all pairs of the cluster transition patterns based on all the cluster transition patterns input by the cluster transition pattern input unit and the centroid distance matrix for all the cluster pairs calculated by the centroid distance matrix calculation unit, and a change point score calculation unit that calculates the distance between the cluster transition tensors of the past period and the current period in consideration of the distance between the cluster transition patterns based on the cluster transition tensors of the past period and the current period input by the cluster transition tensor input unit and the distance matrix for all pairs of the cluster transition patterns calculated by the cluster transition pattern distance matrix calculation unit.
Effect of the Invention
[0011] As described above, according to the present invention, when a cluster transition pattern far from what was observed in the past is newly observed, or when a cluster transition pattern far from what is currently being observed was observed in the past, a mechanism that promotes an increase in the change point score is introduced, and the effect that more precise change point detection can be realized is achieved.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Mode for Carrying Out the Invention
[0013] ●Change Point Detection Device Prerequisite for This Embodiment First, before explaining the change point score calculation device 20 of this embodiment, the change point detection device 10 that is a prerequisite for this embodiment will be explained with reference to FIGS. 1, 2, and 5. Note that the change point detection device 10 and the change point score calculation device 20 are the same device, and are merely given names indicating one aspect of each feature.
[0014] Here, a change point detection device 10 that can detect the occurrence time as a change point when any change occurs in the system state using time series data representing the system state at each time point of a system (S) composed of one or more devices will be explained. Here, the "system state" refers to the operating state of the system represented by quantitative variables such as "number of accesses" and "number of users".
[0015] 〔Functional Configuration〕 First, the functional configuration of the change point detection device 10 will be explained with reference to FIG. 1. FIG. 1 is a diagram showing an example of the functional configuration of the change point detection device.
[0016] As shown in FIG. 1, the change point detection device 10 includes an input unit 11, a time window generation unit 12, a period setting unit 13, a clustering unit 14, a cluster transition series creation unit 15, a cluster transition tensor calculation unit 16, a change point score calculation unit 17, a detection unit 18, and an output unit 19. Note that the "devices" in the "number of devices" and "state of devices" shown below indicate the devices that constitute the system to be detected for change points by the change point detection device 10.
[0017] The input unit 11 inputs time series data representing the system state at each time point of a system (S) composed of one or more devices, and the time series data is composed of data of (number of devices × number of items representing the state of the device) dimensions.
[0018] The time window generation unit 12 divides the time series data input by the input unit 11 with a fixed-length time window, converts the data at each time point from (number of devices × number of items) -dimensional data to (number of devices × number of items × time window length) -dimensional data to generate conversion data, and performs an intermediate output.
[0019] The period setting unit 13 extracts the time series data of a preset past period and current period from the (number of devices × number of items × time window length) -dimensional time series data generated by the time window generation unit 12, and performs an intermediate output.
[0020] The clustering unit 14 classifies the states of the data of (number of devices × number of items × time window length) dimensions at each time point constituting the time series data of the past period and current period extracted by the period setting unit 13 by a clustering method, and performs an intermediate output.
[0021] The cluster transition series creation unit 15 tracks the clusters assigned by the clustering unit 14 for the data of (number of devices × number of items × time window length) dimensions at each time point of the past period and current period in the time axis direction, creates a series of cluster transitions between different clusters for each of the past period and current period, and at the same time assigns the stay period in each cluster to each cluster constituting this cluster transition series, and performs an intermediate output.
[0022] The cluster transition tensor calculation unit 16 extracts the pre-set fixed-length cluster transitions from the cluster transition series created by the cluster transition series creation unit 15, calculates the occurrence probabilities of each cluster transition pattern in the past period and the current period, uses the above-mentioned cluster transition length (the length of the cluster transition) as the rank (i.e., dimension), has the unique values of all the clusters that appeared in the past period and the current period as the indexes of each dimension, calculates the cluster transition tensor with the occurrence probability of the cluster transition pattern as the value for each of the past period and the current period, and performs intermediate output.
[0023] Based on the cluster transition tensors of the past period and the current period calculated by the cluster transition tensor calculation unit 16, the change point score calculation unit 17 calculates the distance between the cluster transition tensor in the past period and the cluster transition tensor in the current period as the degree of change from the past period to the current period, and performs intermediate output.
[0024] When the change point score calculated by the change point score calculation unit 17 exceeds a pre-set threshold, the detection unit 18 detects it as a change point. That is, when the change point score of the system state calculated based on the data (converted data) at each time point exceeds a pre-set threshold, the detection unit 18 detects it as a change point.
[0025] The output unit 19 outputs the change point detected by the detection unit 18.
[0026] 〔Change point detection process〕 Next, the change point detection process (procedure) will be described with reference to FIG. 2. FIG. 2 is a flowchart showing an example of the change point detection process.
[0027] Hereinafter, assuming that the number of devices constituting the system (S) is M, the number of data items representing the system state at each time point is K, and the number of observation time points of the time series data is N, the time series data is composed of N M×K-dimensional data.
[0028] Note that each element of the M×K-dimensional data at each time point is K observed values representing the states of M devices at that time point. Specifically, when the M×K-dimensional data at a certain time point is [x1,···,xK,xK+1,···,x2K,···,x(M-1)K+1,···,xMK], for example, for m = 1,···,M, x(m-1)K+1,···,xmK are the K observed values of the m-th device at that time point.
[0029] Step S11: First, the input unit 11 inputs time-series data composed of N pieces of M×K (number of devices × number of items) - dimensional data. That is, if the M×K-dimensional data at time point n is Xn, the input unit 11 inputs the time-series data {X1,···,XN}.
[0030] Step S12: Next, the time window generation unit 12 divides the time-series data input in Step S11 by a time window of fixed length W, thereby converting the data at each time point from M×K (number of devices × number of items) - dimensional data to M×K×W (number of devices × number of items × time window length) - dimensional data to generate conversion data and perform an intermediate output. Specifically, an M×K×W-dimensional vector Yn = (Xn-(W-1), Xn-(W-2),···, Xn) composed of the M×K-dimensional data Xn-(W-1), Xn-(W-2),···, Xn at time points n-(W-1), n-(W-2),···, n respectively is used as the M×K×W-dimensional data at time point n. Note that when the original M×K-dimensional data Xn is observed for n = 1,···,N, the converted M×K×W-dimensional data Yn can be obtained for n = W,···,N.
[0031] Step S13: Next, the period setting unit 13 extracts the time-series data of a preset past period and current period from the M×K×W (number of devices × number of items × time window length) - dimensional time-series data generated in Step S12. Specifically, when the past period is [s1, e1] and the current period is [s2, e2], the data {Ys1,···, Ye1} of the past period and the data {Ys2,···, Ye2} of the current period are extracted from the M×K×W-dimensional data Yn at time points n = W,···,N.
[0032] Step S14: Next, the clustering unit 14 classifies the state of (e1 - s1 + e2 - s2 + 2) M×K×W (number of devices × number of items × time window length) - dimensional data that constitutes the time - series data of the past period at the time point of length (e1 - s1 + 1) and the current period at the time point of length (e2 - s2 + 1) extracted in step S13 by a clustering method to obtain a cluster series corresponding to the time - series data. Specifically, when the cluster to which the M×K×W - dimensional data Yn at time point n belongs is Cn, a cluster series {Cs1, ···, Ce1} is obtained from the time - series data {Ys1, ···, Ye1} of the past period, and a cluster series {Cs2, ···, Ce2} is obtained from the time - series data {Ys2, ···, Ye2} of the current period. Note that clustering is a process of classifying (e1 - s1 + e2 - s2 + 2) M×K×W - dimensional data into the same cluster for data that are close to each other based on their mutual distances. A cluster series is obtained by arranging the clusters assigned to each M×K×W - dimensional data in time - series order. As the clustering method, a hierarchical method (for example, the single - linkage method, the complete - linkage method, the group - average method, the Ward method, etc.) may be used, or a non - hierarchical method (for example, the K - Means method, etc.) may be used.
[0033] Step S15: Next, the cluster transition series creation unit 15 tracks the clusters assigned in step S14 for the M×K×W (number of devices × number of items × time window length) - dimensional data at each time point in the past period [s1, e1] and the current period [s2, e2] in the time - axis direction, creates a series of cluster transitions between different clusters for each of the past period and the current period, and assigns a stay period in the corresponding cluster to each cluster constituting this cluster transition series. Specifically, taking the cluster series {Cs1, ···, Ce1} obtained from the time - series data {Ys1, ···, Ye1} in the past period [s1, e1] as an example, when the time points τi (i = 1, 2, ···, I) (where τ1 = s1) at which cluster transitions occur between different clusters in the interval [s1, e1] and the cluster after the transition at the time point τi is c(τi), arranging these in time - series order gives a cluster transition series of length I: c(τ1)→c(τ2)→···→c(τI). Also, for each cluster c(τi) constituting this cluster transition series, by assigning the stay period d(τi)=τi + 1−τi (where τI + 1 = e1) in the cluster c(τi), a cluster transition series with stay periods c(τ1)|d(τ1)→c(τ2)|d(τ2)→···→c(τI)|d(τI) is obtained.
[0034] Step S16: Next, the cluster transition tensor calculation unit 16 extracts cluster transitions of a preset fixed length L from the cluster transition series created in step S15, calculates the occurrence probabilities of each cluster transition pattern in the past period and the current period, sets the cluster transition length L as the rank (dimension), has the unique values of all the clusters that appeared in the past period and the current period as the indexes of each dimension, and calculates a cluster transition tensor having the occurrence probabilities of the cluster transition patterns as values for each of the past period and the current period. Specifically, taking as an example the cluster transition series c(τ1) → c(τ2) → ··· → c(τI) of length I obtained from the time series data {Ys1, ···, Ye1} in the past period [s1, e1], (I - (L - 1)) cluster transitions of length L (where L ≤ I) can be extracted from this cluster transition series, and are represented by c(τi-(L - 1)) → c(τi-(L - 2)) → ··· → c(τi) (i = L, ···, I). The cluster transition tensor calculation unit 16 calculates the occurrence probabilities by grouping these (I - (L - 1)) cluster transitions for each pattern, and calculates an L-dimensional cluster transition tensor based on this. Here, the occurrence probability of a cluster transition pattern is a value obtained by dividing the occurrence frequency of the cluster transition pattern by the total occurrence frequency of all the cluster transition patterns.
[0035] Note that as the occurrence frequency of a cluster transition pattern, a value weighted by the residence period of the cluster transition pattern may be used. Hereinafter, for the sake of simplicity, a method of storing the occurrence probabilities of cluster transition patterns in a tensor will be described using an example where L = 2 and the unique values of all the clusters that appeared through the past period and the current period are α, β, and γ. At this time, the cluster transition tensor is two-dimensional, and the indexes of each dimension take three values α, β, and γ. The cluster transition tensor can be represented by a 3×3 array. When the occurrence probability of the cluster transition pattern α → β is 0.1, the occurrence probability 0.1 is stored in the array element where the index of the first axis (the first element of the cluster transition pattern) has the value α and the index of the second axis (the second element of the cluster transition pattern) has the value β.
[0036] Step S17: Next, the change point score calculation unit 17 calculates the degree of change from the past period to the current period as the distance between the cluster transition tensor in the past period and the cluster transition tensor in the current period based on the cluster transition tensors in the past period and the current period calculated in Step S16. When the elements of the cluster transition tensor D1 in the past period are d1i1,···,iL and the elements of the cluster transition tensor D2 in the current period are d2i1,···,iL, the distance between the two can be expressed, for example, by the following mean squared error. (Σl=1LΣm=1M(d2i1,···,iL - d1i1,···,iL)2 / ML)1 / 2 Note that M in the following tensor distance is the number of unique values of all clusters that have appeared through the past period and the current period.
[0037] Step S18: Next, the detection unit 18 detects a change point when the change point score calculated in Step S17 exceeds a preset threshold. That is, the detection unit 18 detects a change point when the change point score of the system state calculated based on the data (converted data) at each time point exceeds a preset threshold.
[0038] Step S19: Finally, the output unit 19 outputs the change point detected in Step S18.
[0039] 〔Hardware Configuration〕 Subsequently, the hardware configuration of the change point detection device 10 will be described with reference to FIG. 3. FIG. 3 is a hardware configuration diagram of the change point detection device.
[0040] As shown in FIG. 3, the change point detection device 10 includes a processor 101, a memory 102, an auxiliary storage device 103, a connection device 104, a communication device 105, and a drive device 106. Each piece of hardware constituting the change point detection device 10 is interconnected via a bus 107.
[0041] The processor 101 serves as a control unit that controls the entire change point detection device 10 and has various arithmetic devices such as a CPU (Central Processing Unit). The processor 101 reads various programs onto the memory 102 and executes them. Note that the processor 101 may include a GPGPU (General-purpose computing on graphics processing units).
[0042] The memory 102 has main memory devices such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The processor 101 and the memory 102 form a so-called computer, and the computer realizes various functions by the processor 101 executing various programs read onto the memory 102.
[0043] The auxiliary storage device 103 stores various programs and various information used when the various programs are executed by the processor 101.
[0044] The connection device 104 is a connection device that connects an external device (for example, a display device 110, an operation device 111) and the change point detection device 10.
[0045] The communication device 105 is a communication device for transmitting and receiving various information to and from other devices.
[0046] The drive device 106 is a device for setting the recording medium 130. The recording medium 130 here includes media that optically, electrically, or magnetically record information, such as a CD-ROM (Compact Disc Read-Only Memory), a flexible disk, and a magneto-optical disk. The recording medium 130 may also include semiconductor memories such as a ROM (Read Only Memory) and a flash memory that electrically record information.
[0047] Note that various programs installed in the auxiliary storage device 103 are installed, for example, when the distributed recording medium 130 is set in the drive device 106 and the various programs recorded on the recording medium 130 are read by the drive device 106. Alternatively, various programs installed in the auxiliary storage device 103 may be installed by being downloaded from a network via the communication device 105.
[0048] 〔Main effects by the change point detection device〕 As described above, the change point detection device 10 can detect the occurrence time as a change point when any change occurs in the system state, using time-series data representing the system state at each time point of a system (S) composed of one or more devices.
[0049] Moreover, since the change point detection device 10 is premised on a method of classifying the system state at each time point by a clustering method, it can target time-series data including data that does not satisfy stationarity constraints or iid constraints such as showing periodic fluctuations. Furthermore, the change point detection device 10 models the periodic fluctuations of the system (S) by considering the state transition of the system (S) over time (that is, the transition of the cluster to which the system state belongs at each time point and its residence period), and can detect changes including changes in time change patterns such as changes in periodic fluctuations.
[0050] Note that the change point score calculation device 20 described later has the same hardware configuration as the change point detection device 10, and thus its description will be omitted.
[0051] ● Change point score calculation device of the present embodiment Next, an embodiment of the present invention will be described. In this embodiment, when the change point detection device 10 calculates the distance between the cluster transition tensors in the past period and the current period in the change point score calculation unit 17, the square error of the residence probability for each cluster transition pattern in the past period and the current period is weighted by the distance from the other patterns of each cluster transition pattern. By doing so, when a cluster transition pattern far from what was observed in the past is newly observed, or when a cluster transition pattern far from what is currently being observed was observed in the past, a mechanism that promotes an increase in the change point score is introduced, and a change point score calculation device 20 that can realize more precise change point detection will be described.
[0052] 〔Functional Configuration〕 First, the functional configuration of the change point score calculation device 20 according to this embodiment will be described with reference to FIG. 3. FIG. 3 is a diagram showing an example of the functional configuration of the change point score calculation device according to this embodiment.
[0053] As shown in FIG. 3, the change point score calculation device 20 according to this embodiment includes a centroid coordinate input unit 21, a cluster transition pattern input unit 22, a cluster transition tensor input unit 23, a centroid distance matrix calculation unit 24, a cluster transition pattern distance matrix calculation unit 25, a change point score calculation unit 26, and an output unit 27.
[0054] The centroid coordinate input unit 21 inputs the centroid (cluster center) coordinates of all the clusters generated by the clustering unit 14 (all the clusters assigned to the data in the dimension of the number of devices × number of items × time window length at each time point constituting the time series data in the past period and the current period).
[0055] The cluster transition pattern input unit 22 inputs all the cluster transition patterns (all the cluster transition patterns that appeared in the past period and the current period) extracted by the cluster transition tensor calculation unit 16.
[0056] The cluster transition tensor input unit 23 inputs the cluster transition tensors for each of the past period and the current period calculated by the cluster transition tensor calculation unit 16.
[0057] The centroid distance matrix calculation unit 24 calculates a centroid distance matrix for all cluster pairs based on the centroid coordinates of all clusters input by the centroid coordinate input unit 21, and performs an intermediate output.
[0058] The distance matrix calculation unit for distances between cluster transition patterns 25 calculates a distance matrix for all cluster transition pattern pairs based on all the cluster transition patterns input by the cluster transition pattern input unit 22 and the centroid distance matrix for all cluster pairs calculated by the centroid distance matrix calculation unit 24, and performs an intermediate output.
[0059] The change point score calculation unit 26 calculates the distance between the cluster transition tensors for the past period and the current period considering the distances between cluster transition patterns based on the cluster transition tensors for each of the past period and the current period input by the cluster transition tensor input unit 23 and the distance matrix for all cluster transition pattern pairs calculated by the distance matrix calculation unit for distances between cluster transition patterns 25, and performs an intermediate output. Note that the change point score calculation unit 26 is a functional unit showing one aspect of the function of the change point score calculation unit 17.
[0060] The output unit 27 outputs the distance between the cluster transition tensors for the past period and the current period calculated by the change point score calculation unit 26 as a change point score, and passes it to the detection unit 18.
[0061] 〔Change point score calculation process〕 Next, the change point score calculation process (procedure) according to the present embodiment will be described with reference to FIG. 4. FIG. 4 is a flowchart showing an example of the change point score calculation process according to the present embodiment.
[0062] Step S21: First, the centroid coordinate input unit 21 inputs the centroid (cluster center) coordinates of all the clusters generated by the clustering unit 14 (all the clusters assigned to the data in the dimension of the number of devices × number of items × time window length at each time point constituting the time series data of the past period and the current period). That is, from the K (number of devices × number of items × time window length)-dimensional time series data constituting the time series of the past period and the current period, if I clusters c1, c2,..., c I are generated by the clustering unit 14, and the centroid coordinates (K-dimensional vector) of the i-th (= 1, 2,..., I) cluster c i are X i =(X i1 , X i2 ,..., X iK ), then the centroid coordinate input unit 21 inputs the I centroid coordinates X1, X2,..., X I .
[0063] Step S22: The cluster transition pattern input unit 22 inputs all the cluster transition patterns (all the cluster transition patterns that appeared in the past period and the current period) extracted by the cluster transition tensor calculation unit 16. That is, if the cluster transition pattern input unit 22 extracts M cluster transition patterns of fixed length L by the cluster transition tensor calculation unit 16, and the m-th (= 1, 2,..., M) cluster transition pattern is {c m1 , c m2 ,..., c mL}, then the cluster transition pattern input unit 22 inputs the M cluster transition patterns {c m1 , c m2 ,..., c mL} (m = 1, 2,..., M).
[0064] Step S13: The cluster transition tensor input unit 23 inputs the cluster transition tensors for each of the past period and the current period calculated by the cluster transition tensor calculation unit 16. That is, if the cluster transition tensors calculated for each of the past period and the current period are D1 and D2 respectively, the cluster transition tensor input unit 23 inputs these two tensors D1 and D2. Here, the cluster transition tensor, both D1 and D2, has the length L of the cluster transition pattern as the rank (number of dimensions), has the unique values of all the clusters that appeared in the past period and the current period as the indices of each dimension, and the stay probability of the cluster transition pattern in which the indices (clusters) of each dimension are arranged in dimension order is stored in the element corresponding to the combination of the indices (clusters) of each dimension.
[0065] Step S24: The centroid distance matrix calculation unit 24 calculates the centroid distance matrix for all cluster pairs based on the centroid coordinates of all the clusters input by the centroid coordinate input unit 21. That is, based on the centroid coordinates X1, X2,..., X I of the I clusters c1, c2,..., c I input by the centroid coordinate input unit 21, for all cluster pairs {c i , c j}(i = 1, 2,..., I, j = 1, 2,..., I), the distance d i , X j} between the centroid coordinate pairs {X c (X i , X j ) is calculated and stored in the (i, j) component to obtain the I-row I-column centroid distance matrix M c . Here, as the distance d i , X j} between the centroid coordinate pairs {X c (X i , X j ), the Euclidean distance represented by the following formula may be used, or other distances (Manhattan distance, Chebyshev distance, Mahalanobis distance, etc.) may be used.
[0066] d c (X i ,X j )=(Σ k=1 K (X jk -X ik ) 2 ) 1 / 2 Also, since the distance d i ,X j} between the centroid coordinate pairs {X c (X i ,X j ) does not change even if clusters i and j are swapped, M c is the target matrix.
[0067] Step S25: Based on the distance matrix between centroids for all the cluster transition patterns input by the cluster transition pattern input unit 22 and all the cluster pairs calculated by the distance matrix calculation unit 24 between centroids, the distance matrix calculation unit 25 for cluster transition patterns calculates the distance matrix for all the cluster transition pattern pairs. That is, for all the cluster transition patterns π m ={c m1 ,c m2 ,...,c mL}(m = 1, 2,..., M) input by the cluster transition pattern input unit 22 and the distances d i ,c j}(i = 1, 2,..., I, j = 1, 2,..., I) between centroids for all the cluster pairs {c c (X i ,X j ) calculated by the distance matrix calculation unit 24 between centroids, the distance matrix calculation unit 25 for cluster transition patterns calculates the distance d m ,π n}(m = 1, 2,..., M, n = 1, 2,..., M) between cluster transition patterns for all the cluster transition pattern pairs {π p (π m ,π n ) and stores this in the (m, n) component to obtain an M-row M-column distance matrix M p .
[0068] Note that for calculating the distance between the cluster transition pattern π m ={c m1 , c m2 ,..., c mL} and the cluster transition pattern π n ={c n1 , c n2 ,..., c nL}, for example, an alignment technique, which is a technique in the field of bioinformatics, is used. The alignment technique is a technique that can be used to determine the similarity between two or more sequences, and is a technique for solving an optimization problem of obtaining an optimal correspondence relationship between sequences while inserting gap symbols so that the sequence lengths are the same. In step S25, in particular, a pairwise alignment technique for obtaining an optimal correspondence relationship between two sequences is used. In the pairwise alignment technique, what is optimized (maximized) is a score calculated from the correspondence relationship between two sequences. For each case where the paired character pair is the same, different, or one is a gap, a score is set in advance, and it is calculated by the sum thereof. In step S25, using this pairwise alignment technique, an optimal correspondence relationship between two cluster transition patterns can be obtained and the maximum score can be calculated. Therefore, this maximum score is regarded as the similarity between two cluster transition patterns, and the distance between two cluster transition patterns can be calculated from this similarity. For example, the distance can be calculated by a method such as subtracting 1 after normalizing the similarity in the interval [0, 1]. Note that in general alignment techniques, a constant score is often set in advance for each case where the paired character pair is the same, different, or one is a gap. However, in the case of step S25, since the characters constituting the sequence (cluster transition pattern) correspond to clusters, the inter-centroid distance calculated in step S24 is used as the score set for different character pairs (cluster pairs). For example, if the score set for the same character pair (cluster pair) is +s and the score set when one is a gap is -g, the inter-centroid distance matrix M cNormalize it so that the minimum value is +s and the maximum value is -g. This is the centroid distance matrix M c Since the minimum value of c corresponds to the distance between the same cluster pair, associate this with the score +s set for the same character pair (cluster pair), and also for the centroid distance matrix M c Since the maximum value of c corresponds to the distance between the most distant cluster pairs, this means associating it with the score -g when one of them is a gap.
[0069] Step S26: The change point score calculation unit 26 calculates the distance between the cluster transition tensors of the past period and the current period considering the distance between the cluster transition patterns, based on the cluster transition tensors of the past period and the current period input by the cluster transition tensor input unit 23 and the distance matrix for all cluster transition pattern pairs calculated by the cluster transition pattern distance matrix calculation unit 25. That is, for the cluster transition tensors D1 and D2 of the past period and the current period input by the cluster transition tensor input unit 23, the residence probabilities of the cluster transition patterns π m ={c m1 ,c m2 ,...,c mL} are p1 m , p2 m , and for the distance matrix M p calculated for all cluster transition pattern pairs by the cluster transition pattern distance matrix calculation unit 25, take the row average of the distance vector
[0070]
Number
[0071] d(D1,D2)=(Σ m δ m (p2 m -p1 m ) 2 ) 1 / 2 Note that, here, the distance vector
[0072]
Number
[0073] δ m = Σ n≠m d p (π m , π n ) / (N - 1) Step S27: The output unit 27 outputs the distance d(D1, D2) between the cluster transition tensors of the past period and the current period calculated by the change point score calculation unit 26 as the change point score and passes it to the detection unit 8.
[0074] 〔Main effects of the embodiment〕 As described above, when calculating the distance between the cluster transition tensors of the past period and the current period in the change point score calculation unit 17 of the change point score calculation device 20 according to the present embodiment, the squared error of the residence probability for each cluster transition pattern in the past period and the current period is weighted by the distance from other patterns of each cluster transition pattern, so that when a cluster transition pattern far from what has been observed in the past is newly observed, or when a cluster transition pattern far from what is currently being observed was observed in the past, a mechanism is introduced to promote an increase in the change point score, and more refined change point detection can be realized.
[0075] 〔Supplementary note〕 The present invention is not limited to the above-described embodiment, and may have the following configurations or processes (operations).
[0076] The change point detection device 10 and the change point score calculation device 20 can also be realized by a computer and a program. However, it is also possible to record this program on a (non-transitory) recording medium or provide it through a network such as the Internet.
Explanation of Signs
[0077] 10 Change point detection device 11 Input unit 12 Time window generation unit 13 Period setting unit 14 Clustering unit 15 Cluster transition series creation unit 16 Cluster transition tensor calculation unit 17 Change point score calculation unit 18 Detection unit 19 Output unit 20 Change point score calculation device 21 Centroid coordinate input unit 22 Cluster transition pattern input unit 23 Cluster transition tensor input unit 24 Centroid distance matrix calculation unit 25 Cluster transition pattern distance matrix calculation unit 26 Change point score calculation unit 27 Output unit
Claims
1. A centroid coordinate input unit that inputs centroid coordinates that are the centers of all clusters assigned to data of the dimension of the number of devices × number of items × time window length at each time point constituting the time series data of the past period and the current period generated by a clustering unit; A cluster transition pattern input unit that inputs all cluster transition patterns that appeared in the past period and the current period, which were extracted in a cluster transition tensor calculation unit; A cluster transition tensor input unit that inputs the cluster transition tensors for each of the past period and the current period, which were calculated in the cluster transition tensor calculation unit; A centroid distance matrix calculation unit that calculates a centroid distance matrix for all cluster pairs based on the centroid coordinates of all clusters input by the centroid coordinate input unit; A cluster transition pattern distance matrix calculation unit that calculates a distance matrix for all pairs of the cluster transition patterns based on all the cluster transition patterns input by the cluster transition pattern input unit and the centroid distance matrix for all the cluster pairs calculated by the centroid distance matrix calculation unit; A change point score calculation unit that calculates the distance between the cluster transition tensors of the past period and the current period in consideration of the distance between the cluster transition patterns, based on the cluster transition tensors for each of the past period and the current period input by the cluster transition tensor input unit and the distance matrix for all pairs of the cluster transition patterns calculated by the cluster transition pattern distance matrix calculation unit; A change point score calculation device having the above components.
2. The cluster transition pattern distance matrix calculation unit uses an alignment technique when calculating the distance between the cluster transition patterns, regards a value obtained by normalizing a score for achieving an optimal correspondence relationship between the cluster transition patterns calculated by the alignment technique as a similarity, and calculates the distance between the cluster transition patterns based on the similarity. The change point score calculation device according to Claim 1.
3. A centroid coordinate input process for inputting centroid coordinates that are the centers of all clusters assigned to data of the dimension of the number of devices × number of items × time window length at each time point constituting the time series data of the past period and the current period generated by a clustering unit; A cluster transition pattern input process for inputting all cluster transition patterns that appeared in the past period and the current period, which were extracted in the cluster transition tensor calculation unit, and A cluster transition tensor input process for inputting the cluster transition tensors for the past period and the current period, respectively, which were calculated in the cluster transition tensor calculation unit, and A centroid distance matrix calculation process for calculating a centroid distance matrix for all cluster pairs based on the centroid coordinates of all clusters input by the centroid coordinate input process, and A cluster transition pattern distance matrix calculation process for calculating a distance matrix for all pairs of the cluster transition patterns based on all the cluster transition patterns input by the cluster transition pattern input process and the centroid distance matrix for all the cluster pairs calculated by the centroid distance matrix calculation process, and A change point score calculation process for calculating the distance between the cluster transition tensors for the past period and the current period in consideration of the distance between the cluster transition patterns based on the cluster transition tensors for the past period and the current period, respectively, input by the cluster transition tensor input process and the distance matrix for all pairs of the cluster transition patterns calculated by the cluster transition pattern distance matrix calculation process, and A change point score calculation method executed by a computer.
4. A program for causing a computer to execute the method according to claim 3.
Citation Information
Patent Citations
Clustering program, clustering method, and information processing apparatus
JP2017068748A