A pattern discovery method, system and terminal for unlabeled multidimensional time series data

By calculating the clustering labels from each dimension perspective and converting them into undirected weighted graphs, community discovery processing is carried out, and the problems of poor clustering effect and slow speed of multi-dimensional time series data are solved, efficient and accurate clustering results are achieved, and artificial interference is reduced.

CN114611620BActive Publication Date: 2025-05-23SOUTHWEST PETROLEUM UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210265902.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-05-23
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

The prior art fails to fully consider the impact of each dimension attribute on the clustering results when clustering multi-dimensional time series data, resulting in poor clustering effect. At the same time, the clustering speed is slow and the number of tags needs to be manually specified, which increases manual interference.

Method used

By calculating the clustering labels of multidimensional timing data from each dimension perspective, it is converted into a collection of correlation matrixes, and merged into a multidimensional attribute feature information similarity matrix to form an undirected weighted graph, and community discovery processing is carried out to obtain the pattern of multidimensional timing data.

Benefits of technology

It improves clustering accuracy, reduces manual interference, improves clustering speed and efficiency, and reduces human and financial costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611620B_ABST
    Figure CN114611620B_ABST
Patent Text Reader

Abstract

The present invention discloses a pattern discovery method, system and terminal for unlabeled multidimensional time series data, belonging to the field of clustering technology, and the method includes: calculating the clustering labels of multidimensional time series data from the perspective of each dimension and converting them into a correlation matrix set, merging the sets into a multidimensional attribute feature information similarity matrix, and converting them into an undirected weighted graph; performing community discovery processing based on the undirected weighted graph to obtain the pattern of the multidimensional time series data. The present invention calculates the clustering labels of multidimensional time series data from the perspective of each dimension, taking into account the similarity between the attributes of each dimension; based on this, a multidimensional attribute feature information similarity matrix containing information of each dimension is obtained, which fully considers the influence of dimensional information on pattern discovery results, thereby improving clustering accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of clustering technology, and in particular to a pattern discovery method, system and terminal for unlabeled multidimensional time series data. Background Art

[0002] With the development of computer technology, data in various fields can be stored in the form of time series. Clustering time series data for pattern discovery has been applied to different industries, and these patterns enable data analysts to extract valuable information from complex and large-scale data sets.

[0003] Time series data is divided into univariate time series data and multidimensional time series data according to its attribute dimension. In the real world, most of the collected and stored data is multidimensional time series data. This type of data has become a more complex data type in the field of data analysis due to its long time dimension and many attribute variables. In addition, since most of the time series data collected and stored in the real world is unlabeled data, if the supervised method in mainstream machine learning is used for data analysis, such data needs to be manually labeled, resulting in a waste of human resources and low efficiency.

[0004] Therefore, using unsupervised methods to analyze and discover patterns in multidimensional time series data can reduce time and labor costs and improve efficiency. Due to the high-dimensional and complex characteristics of multidimensional time series data, there are relatively few research results in related areas. At present, the main problems in the research on multidimensional time series data clustering are:

[0005] 1. In multidimensional time series data, the data of each attribute dimension has a significant impact on the clustering results and discovered patterns.

[0006] 2. Due to the large volume of time series data, the time series similarity measurement and clustering speed are slow, especially when considering multidimensional time series data with multiple dimensional attributes, the efficiency is even lower.

[0007] 3. Some clustering algorithms require manual input of the number of clustering labels, which increases the human interference in the pattern discovery results. Summary of the invention

[0008] The purpose of the present invention is to solve the problem that the existing technology does not consider the influence of multi-dimensional attributes on clustering results when discovering patterns in multi-dimensional time series data, resulting in poor clustering effect, and provides a pattern discovery method, system and terminal for unlabeled multi-dimensional time series data.

[0009] The objective of the present invention is achieved through the following technical solution: a pattern discovery method for unlabeled multidimensional time series data, the method comprising the following steps:

[0010] Calculate cluster labels for multidimensional time series data from each dimensional perspective And transformed into a correlation matrix set

[0011] Will gather Merge into a multi-dimensional attribute feature information similarity matrix and transform into an undirected weighted graph;

[0012] The community discovery process based on undirected weighted graphs is used to obtain patterns of multi-dimensional time series data.

[0013] In one example, the clustering labels of multidimensional time series data under each dimensional perspective are calculated. Specifically include:

[0014] Extract the component data of each dimension of the multidimensional time series data and select the initial vector center;

[0015] Calculate the distance difference between each component feature vector and the center of the initial vector to obtain the preliminary clustering results;

[0016] The preliminary clustering results are subjected to clustering iteration processing. During the clustering iteration process, the distance difference between the feature component and the initial vector center is calculated, and the minimum distance difference is calculated to obtain the optimal clustering vector center, and then the optimal component data clustering result is obtained.

[0017] In one example, selecting the initial vector center specifically includes:

[0018] The component data are symmetrically split, and the influencing factors of each component in the multi-dimensional data are summed and averaged to obtain vector data distributed in two-dimensional space, and then the initial vector center is selected.

[0019] In one example, the clustering iteration process includes:

[0020] Set the number of iterations based on the distribution characteristics and data distribution of multidimensional time series data.

[0021] In one example, performing clustering iteration processing on the preliminary clustering results specifically includes:

[0022] Perform preliminary clustering of feature components based on the initially selected vector center, and conduct preliminary clustering conclusion analysis in a two-dimensional plane;

[0023] The multi-dimensional feature components are divided into two-dimensional feature vectors after absolute value sum and average calculation, and clustered using the k-means method to obtain the cluster standard center;

[0024] Iterate the generated cluster standard center to obtain the clustering results of all the two-dimensional components.

[0025] In one example, the mode of obtaining multi-dimensional time series data by performing community discovery processing based on an undirected weighted graph specifically includes:

[0026] S31: Initialize each vertex of the undirected weighted graph as a community;

[0027] S32: Merge each vertex with its adjacent vertices in turn, calculate the modularity gain ΔQ, and then update the vertices in the community according to the modularity gain ΔQ;

[0028] S33: iterate step S32 until the algorithm is stable;

[0029] S34: compress all nodes in each community into one node, convert the weights of points in the community into the weights of the new node ring, and convert the community construction weights into the weights of the new node edges;

[0030] S35: Repeat steps S31-S33 until the algorithm is stable and the pattern of multi-dimensional time series data is obtained.

[0031] In one example, updating the vertices in the community according to the modularity gain ΔQ specifically includes:

[0032] If the modularity gain ΔQ is greater than 0, the current node is placed in the community where the adjacent node is located.

[0033] It should be further explained that the technical features corresponding to the above examples can be combined or replaced with each other to form a new technical solution.

[0034] The present invention also includes a clustering system for unlabeled multidimensional time series data, the system comprising:

[0035] Multidimensional attribute relationship matrix generation module, used to calculate the clustering labels of multidimensional time series data from each dimensional perspective And transformed into a correlation matrix set

[0036] Multi-dimensional attribute similarity network building module is used to transform the set Merge into a multi-dimensional attribute feature information similarity matrix and transform into an undirected weighted graph;

[0037] The pattern discovery module is used to perform community discovery processing based on an undirected weighted graph to obtain patterns of multi-dimensional time series data.

[0038] In one example, the system further includes a data reading module for converting input multi-dimensional time series data into a matrix.

[0039] The present invention also includes a terminal, including a memory and a processor, wherein the memory stores computer instructions that can be run on the processor, and is characterized in that: when the processor runs the computer instructions, it executes any one of the above examples or a combination of multiple examples to form the steps of the pattern discovery method for unlabeled multidimensional time series data.

[0040] The present invention also includes a storage medium on which computer instructions are stored. When the computer instructions are executed, the steps of the pattern discovery method for unlabeled multidimensional time series data formed by any one or more of the above examples are executed.

[0041] The present invention also includes a terminal, including a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and when the processor runs the computer instructions, it executes the steps of the pattern discovery method for unlabeled multidimensional time series data formed by any one or more of the above examples.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] 1. In one example, by calculating the clustering labels of multidimensional time series data from the perspective of each dimension, the similarity between the attributes of each dimension is taken into account; based on this, a multidimensional attribute feature information similarity matrix containing information of each dimension is obtained, which fully considers the impact of dimensional information on pattern discovery results, thereby improving clustering accuracy.

[0044] 2. In one example, community discovery is performed based on an undirected weighted graph of the multidimensional attribute feature information similarity matrix to obtain the clustering pattern of the multidimensional time series data. There is no need to manually specify the number of patterns of the multidimensional time series data, which reduces manual interference in the pattern discovery results. At the same time, it can improve the speed and efficiency of the traditional multidimensional time series data clustering algorithm, and greatly reduce the manpower and financial costs compared to manual labeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings. The accompanying drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The same reference numerals are used in these drawings to represent the same or similar parts. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.

[0046] Figure 1 A method flow chart in an example of the present invention;

[0047] Figure 2 The present invention is a method flow chart of a preferred example. DETAILED DESCRIPTION

[0048] The technical solution of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] In the description of the present invention, it should be noted that the directions or positional relationships indicated by "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc. are directions or positional relationships based on the drawings, which are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0050] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, "installation", "connection" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0051] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0052] The specific implementation part of the present invention specifically takes the Lp1 data set in the Robot execution failure data set in industrial data as an example to illustrate the inventive concept of the present application. There are 88 multi-dimensional time series data in the data set, and each time series data has 6 dimensional attributes.

[0053] In one example, if Figure 1 As shown, a pattern discovery method for unlabeled multidimensional time series data specifically includes the following steps:

[0054] S1: Calculate the cluster labels of multidimensional time series data from each dimension perspective And transformed into a correlation matrix set Among them, the cluster label is used to mark the data mode to which the current dimension time series data belongs.

[0055] S2: Set Merge into a multi-dimensional attribute feature information similarity matrix and transform into an undirected weighted graph;

[0056] S3: Perform community discovery based on the undirected weighted graph to obtain the pattern of the multidimensional time series data. In this application, the pattern is the data category to which the multidimensional time series data belongs; and pattern discovery is used to determine the data category to which the multidimensional time series data belongs.

[0057] This application calculates the clustering labels of multidimensional time series data from the perspective of each dimension, takes into account the similarity between the attributes of each dimension, and on this basis, clusters the multidimensional time series data as a whole based on the multidimensional attribute feature information similarity matrix containing the information of each dimension. That is, in the overall clustering process of multidimensional time series data, the influence of dimensional information on pattern discovery results is fully considered, thereby improving the clustering accuracy and obtaining a clustering result that fits the actual data distribution.

[0058] In one example, cluster labels for multidimensional time series data are calculated from each dimension perspective. Specifically include:

[0059] S11: extracting component data of each dimension of the multidimensional time series data, performing partitioning and averaging processing, and selecting an initial vector center to perform clustering processing on each component data;

[0060] S12: Calculate the distance difference between each component feature vector and the vector center to divide the area to which each sample belongs, that is, to achieve preliminary clustering processing. Specifically, according to the two-dimensional initial vector center obtained by the partition and average processing of the multidimensional data in S11, calculate the distance difference between the two-dimensional data points (components) processed by dimensionality reduction in LP1 and the initial center point, and compare the distance difference of 88 LP1 data with the distance difference of the initial vector center to obtain preliminary partition clustering.

[0061] S13: Continue to iterate on the result of the preliminary clustering to determine whether the distance difference between the component feature and the center of the initial vector has reached the extreme. If the distance difference is abnormal after continuing the iteration, the last distance difference is taken as the critical value. The clustering result at this time is the final clustering result, i.e., the component data clustering result. Stop iteration.

[0062] In one example, selecting the initial vector center specifically includes:

[0063] The component data is symmetrically split, and the influencing factors of each component in the multidimensional data are summed and averaged to obtain vector data distributed in two-dimensional space, which reduces the amount of data processing for subsequent clustering calculations. On this basis, the initial vector center is further selected; specifically, for the multivariate time series A=[A 1 ,A 2 ,…,A m ] and multivariate time series B = [B 1 ,B2 ,…,B m ], split the data in the component symmetrically, calculate the absolute value and average:

[0064]

[0065] a n =|A 1 +A 2 +…A v | / v

[0066] b n =|B 1 +B 2 +…B v | / v

[0067] Where v represents the boundary where the multidimensional time series data is divided according to the number of component attributes; m represents the number of component data in the multivariate time series; a n represents the sum of the absolute values ​​of each component data in the multivariate time series A; b n represents the sum of the absolute values ​​of each component data in the multivariate time series B. On this basis, the vector data data (a n ,b n ) and construct a two-dimensional space to provide a visual selection framework for the selection of vector centers.

[0068] Specifically, step S12 calculates the distance difference between each component feature vector and the vector center, which is:

[0069] According to the distribution of component data in each area, the distance between the vector centers of the acquired component two-dimensional space transformation data is calculated to calculate the distribution distance D of the divided intervals:

[0070]

[0071] Among them, data represents the eigenvector of the component data; centerPoint(labels) represents the center vector. By calculating the distance support partition interval selection, apply the sample point data(a n ,b n ) divides the area to which the sample belongs, and obtains the similarity matrix Y of the local relationship of the component attribute sequence = {Y 1 ,Y 2 ,…,Y 6}.

[0072] In one example, it also includes updating the initial vector center based on the component tolerance change to obtain the optimal vector center. Specifically, after obtaining the distance of the logical partition interval, the data of each dimension contained in the multidimensional time series data set is extracted separately, the data features of each dimension are compared, and the relevant single-dimensional data is selected for processing, and the change of the component tolerance Var |newVar-oldVar|≥toal is judged, where oldVar represents the component tolerance obtained by the last clustering process; newVar represents the current component tolerance; when it is less than the total cumulative tolerance toal, it is selected as the initial vector center; according to the delineated initial vector center, the formula for calculating the distance matrix dist from data to the initial clustering center centerPoint is:

[0073]

[0074] Where T represents transpose; the distance between each point in the matrix and the center point is dist[i][:], which represents the distance from the i points to the generated n centers.

[0075] In one example, the clustering iteration process includes:

[0076] The number of iterations is set according to the distribution characteristics and data distribution of the multidimensional time series data. Specifically, the number of iterations is determined by the size of the multidimensional time series data set. Too many iterations will lead to overfitting and cause the vector center to be out of order. The number of iterations of the algorithm is set according to the distribution characteristics and data distribution of the data set. The calculation method is: data-centerPoint(labels) 2 , count+1, and finally returns the number of iterations count. The iterative method using the function descent principle is used. Each repetition of the calculation process is called an "iteration", and the result of each iteration will be used as the initial value of the next iteration, and finally the optimal clustering result is obtained. Among them, the iterative calculation using the function descent principle is publicized as:

[0077] |f(X (k+1) )-f(X k )|≤ε,(|f(X (k+1) )|≤1),

[0078] Among them, f(X k ) represents the current iteration sequence; f(X (k+1) ) represents the current next iteration sequence; ε represents the error threshold;

[0079] In one example, performing clustering iteration processing on the preliminary clustering results specifically includes:

[0080] S141: converting the component regression calculation into a space vector according to the vector center and the number of iterations;

[0081] S142: performing two-dimensional division on the spatial node information after the component tangent sum and average, and obtaining the cluster standard center;

[0082] S143: Perform iterative clustering processing on the generated clustering standard center. During the iterative clustering processing, the distance difference between the characteristic component and the initial vector center is calculated. During the iterative process, the vector center is continuously updated according to the obtained distance difference, and finally the optimal clustering vector center is obtained. The iteration ends, and the component data clustering result is obtained. Specifically, according to the obtained vector center and the optimal number of iterations, the component regression calculation is converted into a spatial vector, and then the spatial node information after the sum and average of these components is divided into two dimensions to obtain the clustering standard center; the specific division method is:

[0083]

[0084] Among them, θ represents the angle of the eigenvector in the two-dimensional coordinate system; a represents the initial center horizontal coordinate; b represents the initial center vertical coordinate. Then the number of centers obtained is verified by the elbow method. The verification method is:

[0085]

[0086] Among them, SSE represents the clustering error of all samples, indicating the quality of clustering effect; x represents the sample point after the data in LP1 is processed; μ i Represents the centroid of each cluster (the mean of all samples in the initial cluster). Finally, the generated index centers are clustered using the improved partitioning underlying clustering algorithm, and count iterations are performed to finally obtain the specific clustering results of the multidimensional time series data set, return the component clustering labels, and obtain the classification of the similarity matrix, that is, Then we can obtain the multidimensional time series data clustering results from each dimension of the data set.

[0087] Furthermore, step S1 clusters the results Transformed into a correlation matrix set Specifically:

[0088] In the clustering results, the data objects that are classified into the same category are regarded as 1, and the data objects of different categories are regarded as 0. Transformed into a correlation matrix that reflects the multidimensional time series data objects from different dimensional perspectives In this embodiment, It is a 88×88 matrix.

[0089] Furthermore, the collection Merge into a multi-dimensional attribute feature information similarity matrix, the merging formula is In this embodiment in, Represents a collection of correlation matrices under a single dimension.

[0090] Furthermore, the multi-dimensional attribute feature information similarity matrix is ​​transformed into an undirected weighted graph G, G = <V L ,E L >. Among them, V L represents a node set in an undirected weighted graph. In this embodiment, there are 88 nodes, corresponding to 88 multi-dimensional time series data, that is, each multi-dimensional time series data in the matrix is ​​initialized as a node in the graph. L Represents the edge set ES= <V i ,weight,V j >, where the value of weight is a matrix Zhongyu The corresponding eigenvalues ​​are Represents the set of correlation matrices under a single dimension attribute j. and The corresponding eigenvalue is initialized to the vertex V in the graph i With V j The value of the connected edges is used to associate the component data of each dimension with the undirected weighted graph. On the basis of fully considering the impact of the component data on the overall pattern clustering results of the multidimensional time series data, the multidimensional time series data is converted into an undirected weighted graph, and then the community discovery algorithm is introduced to cluster the multidimensional time series data as a whole again. While ensuring the clustering accuracy, the clustering time cost of the multidimensional time series data is greatly reduced.

[0091] In one example, the mode of obtaining multi-dimensional time series data by performing community discovery processing based on an undirected weighted graph specifically includes:

[0092] S31: Initialize each vertex of the undirected weighted graph as a community. Here, the vertex represents multidimensional time series data, and the community represents the clustering pattern. In this example, the initial number of communities is 88.

[0093] S32: Merge each vertex with its adjacent vertex in turn, calculate the modularity gain ΔQ between the two, and then update the vertex information in the community according to the modularity gain ΔQ;

[0094] S33: Iterate step S2 until the algorithm is stable, that is, the communities to which all vertices belong no longer change.

[0095] S34: compress all nodes (vertices) of each community into one node, convert the weights of points within the community into the weights of the new node ring, and convert the community construction weights into the weights of the new node edges;

[0096] S35: Repeat steps S31-S33 until the algorithm is stable, the pattern of the multidimensional time series data is obtained, and the multidimensional time series data is divided into different patterns.

[0097] Specifically, the calculation formula of the modularity gain ΔQ in step S32 is:

[0098]

[0099] Among them, m is the sum of all weighted degrees in the entire graph; K i Represents the sum of the weights of the edges connecting node i and all the nodes in the undirected weighted graph; if ΔQ>0, the node is placed in the community where the adjacent node is located.

[0100] In this embodiment, by performing pattern discovery, i.e., data clustering, on the Lp1 data set, different error patterns of robot execution errors in industrial data are obtained. The present invention can be applied to pattern discovery of multi-dimensional time series data collected by industrial sensors.

[0101] Combining the above examples, we can obtain the preferred examples of this application, such as Figure 2 As shown, the specific steps include:

[0102] S1': extract the component data of each dimension of the multidimensional time series data, perform partitioning and averaging, and select the initial vector center;

[0103] S2': Calculate the distance difference between each component feature vector and the vector center, and perform preliminary clustering processing;

[0104] S3': Perform clustering iteration on the preliminary clustering results. During the clustering iteration process, the distance difference between the feature component and the initial vector center is calculated, and the minimum distance difference is calculated to obtain the optimal clustering vector center, and then the optimal component data clustering result is obtained.

[0105] S4': Clustering results Transformed into a set of correlation matrices

[0106] S5': The collection Merge into a multi-dimensional attribute feature information similarity matrix;

[0107] S6': Convert the multi-dimensional attribute feature information similarity matrix into an undirected weighted graph;

[0108] S7': Initialize each vertex of the undirected weighted graph as a community;

[0109] S8': Merge each vertex with its adjacent vertices in turn, and calculate the modularity gain ΔQ of the two. Then, update the vertex information in the community according to the modularity gain ΔQ, and iterate until the algorithm is stable.

[0110] S9': compress all nodes in each community into one node, convert the weights of points in the community into the weights of the new node ring, and convert the community construction weights into the weights of the new node edges;

[0111] S10': Repeat steps S8'-S9' until the algorithm is stable and the pattern of the multi-dimensional time series data is obtained.

[0112] The present invention also includes a clustering system for unlabeled multidimensional time series data, the system comprising:

[0113] Multidimensional attribute relationship matrix generation module, used to calculate the clustering labels of multidimensional time series data from each dimensional perspective And transformed into a correlation matrix set

[0114] Multi-dimensional attribute similarity network building module is used to transform the set Merge into a multi-dimensional attribute feature information similarity matrix and transform into an undirected weighted graph;

[0115] The pattern discovery module is used to perform community discovery based on undirected weighted graphs to obtain patterns of multi-dimensional time series data. Transformed into a set of correlation matrices And merged into a multi-dimensional attribute feature information similarity matrix

[0116] The system of the present invention also includes a data reading module, which is used to convert the input multi-dimensional time series data into a matrix.

[0117] The present application also includes a storage medium having the same inventive concept as Example 1, on which computer instructions are stored, and when the computer instructions are executed, the steps of the above-mentioned pattern discovery method for unlabeled multidimensional time series data are executed.

[0118] Based on this understanding, the technical solution of this embodiment, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0119] The present application also includes a terminal, which has the same inventive concept as that of Example 1, including a memory and a processor, wherein the memory stores computer instructions that can be run on the processor, and the processor executes the steps of the above-mentioned pattern discovery method for unlabeled multidimensional time series data when running the computer instructions. The processor can be a single-core or multi-core central processing unit or a specific integrated circuit, or one or more integrated circuits configured to implement the present invention.

[0120] Each functional unit in the embodiment provided by the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0121] The above specific implementation methods are detailed descriptions of the present invention. It cannot be determined that the specific implementation methods of the present invention are limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions and substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A pattern discovery method for unlabeled multidimensional time series data. Features: It includes the following steps: Calculate cluster labels for multidimensional time series data from each dimensional perspective And transformed into a correlation matrix set The multi-dimensional time series data is the data in the Lp1 dataset in the Robot execution failure dataset in the industrial data; Will gather Merge into a multi-dimensional attribute feature information similarity matrix And transformed into an undirected weighted graph G, G = <V L ,E L > V L Represents a node set in an undirected weighted graph, the matrix Each multidimensional time series data is initialized as a node in an undirected weighted graph; L Represents the edge set ES= <V i ,weight,V j >, weight value is a matrix Zhongyu The corresponding eigenvalues ​​are Represents the set of correlation matrices under a single dimension attribute j. and The corresponding eigenvalue is initialized to the vertex V in the undirected weighted graph i With V j The value of the connected edge, Represents the set of correlation matrices under a single dimension attribute i; The pattern of multidimensional time series data is obtained by community discovery based on undirected weighted graphs. The pattern of multidimensional time series data is the category to which the data in the Lp1 dataset belongs, specifically the different error patterns of robot execution errors in industrial data. The mode of obtaining multi-dimensional time series data by performing community discovery processing based on an undirected weighted graph specifically includes: S31: Initialize each vertex of the undirected weighted graph as a community; S32: Merge each vertex with its adjacent vertices in turn, calculate the modularity gain ΔQ, and then update the vertices in the community according to the modularity gain ΔQ; S33: iterate step S32 until the algorithm is stable; S34: compress all nodes in each community into one node, convert the weights of points in the community into the weights of the new node ring, and convert the community construction weights into the weights of the new node edges; S35: Repeat steps S31-S33 until the algorithm is stable and the pattern of multi-dimensional time series data is obtained.

2. According to the method for pattern discovery of unlabeled multidimensional time series data in claim 1, Features: The clustering labels C of multidimensional time series data under each dimension perspective are calculated. Ym Specifically include: Extract the component data of multidimensional time series data in each dimension and select the initial vector center; Calculate the distance difference between each component feature vector and the center of the initial vector to obtain the preliminary clustering results; The preliminary clustering results are subjected to clustering iteration processing. During the clustering iteration process, the distance difference between the feature component and the initial vector center is calculated, and the minimum distance difference is calculated to obtain the optimal clustering vector center, and then the optimal component data clustering result C is obtained. Ym .

3. According to claim 2, a pattern discovery method for unlabeled multidimensional time series data, Features: The selecting of the initial vector center specifically includes: The component data are symmetrically split, and the influencing factors of each component in the multi-dimensional data are summed and averaged to obtain vector data distributed in two-dimensional space, and then the initial vector center is selected.

4. According to claim 2, a pattern discovery method for unlabeled multidimensional time series data, Features: The clustering iterative process includes: Set the number of iterations based on the distribution characteristics and data distribution of multidimensional time series data.

5. According to claim 2, a pattern discovery method for unlabeled multidimensional time series data, Features: The clustering iteration process for the preliminary clustering results specifically includes: Perform preliminary clustering of feature components based on the initially selected vector center, and conduct preliminary clustering conclusion analysis in a two-dimensional plane; The multi-dimensional feature components are divided into two-dimensional feature vectors after absolute value sum and average calculation, and clustered using the k-means method to obtain the cluster standard center; Iterate the generated cluster standard center to obtain the clustering results C of all the two-dimensional components. Ym .

6. According to the method for pattern discovery of unlabeled multidimensional time series data in claim 1, Features: The updating of the vertices in the community according to the modularity gain ΔQ specifically includes: If the modularity gain ΔQ is greater than 0, the current node is placed in the community where the adjacent node is located.

7. A pattern discovery system for unlabeled multidimensional time series data, Features: It includes: Multidimensional attribute relationship matrix generation module, used to calculate the clustering labels of multidimensional time series data from each dimensional perspective And transformed into a correlation matrix set The multi-dimensional time series data is the data in the Lp1 dataset in the Robot execution failure dataset in the industrial data; Multi-dimensional attribute similarity network building module is used to transform the set Merge into a multi-dimensional attribute feature information similarity matrix and transform into an undirected weighted graph G, G = <V L ,E L > V L Represents a node set in an undirected weighted graph, the matrix Each multidimensional time series data is initialized as a node in an undirected weighted graph; L Represents the edge set ES= <V i ,weight,V j >, weight value is a matrix Zhongyu The corresponding eigenvalues ​​are Represents the set of correlation matrices under a single dimension attribute j. and The corresponding eigenvalue is initialized to the vertex V in the undirected weighted graph i With V j The value of the connected edge, Represents the set of correlation matrices under a single dimension attribute i; The pattern discovery module is used to perform community discovery processing based on an undirected weighted graph to obtain the pattern of multidimensional time series data. The pattern of the multidimensional time series data is the category to which the data in the Lp1 data set belongs, specifically, different error patterns of robot execution errors in industrial data; The mode of obtaining multi-dimensional time series data by performing community discovery processing based on an undirected weighted graph specifically includes: S31: Initialize each vertex of the undirected weighted graph as a community; S32: Merge each vertex with its adjacent vertices in turn, calculate the modularity gain ΔQ, and then update the vertices in the community according to the modularity gain ΔQ; S33: iterate step S32 until the algorithm is stable; S34: compress all nodes in each community into one node, convert the weights of points in the community into the weights of the new node ring, and convert the community construction weights into the weights of the new node edges; S35: Repeat steps S31-S33 until the algorithm is stable and the pattern of multi-dimensional time series data is obtained.

8. A pattern discovery system for unlabeled multidimensional time series data according to claim 7, Features: The system also includes a data reading module, which is used to convert the input multi-dimensional time series data into a matrix.

9. A terminal comprising a memory and a processor, wherein the memory stores computer instructions executable on the processor, Features: When the processor runs the computer instructions, it performs the steps of a pattern discovery method for unlabeled multidimensional time series data as described in any one of claims 1-6.

Citation Information

Patent Citations

  • K-Means clustering lane flow analysis method based on a Gaussian regression algorithm

    CN109800801A

  • Reference subset selection method and system based on consistency clustering and storage medium

    CN113077011A