Asymmetric time series data depth clustering method based on multi-feature fusion

Through the asymmetric time-series data depth clustering method of multi-feature fusion, time-series data features are extracted using recursive graphs, CNNs and LSTMs, combined with the variational autoencoder structure, the problems of noise and feature extraction in time-series data clustering are solved, and better clustering effect is achieved.

CN120296440APending Publication Date: 2025-07-11INNER MONGOLIA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311297769.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-09
Publication Date
2025-07-11

Smart Images

  • Figure CN120296440A_ABST
    Figure CN120296440A_ABST
Patent Text Reader

Abstract

The invention provides a multi-feature fusion asymmetric time series data depth clustering method, which comprises the following steps of: inputting time series data into an encoder to obtain a first feature, converting the time series data into a recursion plot RP, inputting the recursion plot RP into a convolutional layer, converting the dimension of an output value into the dimension which is the same as the dimension of the first feature to obtain a second feature; fusing the first feature and the second feature through a fusion layer; sampling a fusion layer to obtain a mean value and a variance, randomly sampling from N (0, 1) to obtain epsilon, obtaining embedded layer data, constructing a loss function of an embedded module, predicting the probability that each sample belongs to each cluster, obtaining model prediction, and constructing a target variable based on a prediction model; empirical distribution of the target variables is defined, and KL divergence between the empirical distribution and uniform distribution is minimized; constructing a loss function of a regularization clustering layer, and combining the loss function of the embedded module to obtain a model loss function; training a clustering model based on the loss function and the target variable; and constructing a clustering model based on the new target variable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a deep clustering method, and in particular to a deep clustering method for asymmetric time series data with multi-feature fusion. Background Art

[0002] Clustering is an unsupervised learning method that divides the samples in a dataset into several clusters, each cluster consisting of usually disjoint subsets, where the objects within a cluster have the greatest similarity and the objects between clusters have the greatest dissimilarity. When processing time series data, its label information usually represents certain specific states, events, behaviors, etc. These information play a key role in tasks such as pattern discovery, anomaly detection, and information retrieval. For example, a video website can cluster user browsing data to divide different user groups to achieve personalized content push, recommend interesting content to users, and increase user stickiness. In industrial production, clustering temperature, light, vibration and other data can find out whether the machine is abnormal and take targeted preventive measures. Sensors and devices in the Internet of Things generate a large amount of time series data, such as temperature, humidity, air pressure, smart home sensors, etc. Clustering these data can better understand the device behavior and trends in the Internet of Things. Therefore, clustering time series data is of great significance in practical applications.

[0003] Time series data is also affected by high-dimensional high-order data, complex features, and noise. When using expression as the feature extraction link for clustering, the features after expression are input into the clustering algorithm to obtain the clustering result. At this time, although the clustering algorithm does not need to pay attention to the data noise problem, and the features after representation are clearer than the original data, the clustering algorithm still needs the ability to recognize patterns such as periodicity and trendiness to obtain good results. Therefore, even if traditional clustering algorithms use high-quality expressed data as the dataset, the improvement of clustering quality is limited. Currently, existing time series clustering algorithms, such as the shape-based time series data clustering algorithm SBC (Shape-Based Clustering), the spectral-based clustering algorithm SCTSD (Spectral Clustering for Time Series Data), etc., can achieve good clustering results in specific scenarios, but they cannot simultaneously handle the problems of high-dimensional high-order data, complex features, and noise in time series data. In addition, deep learning-based methods have the advantage of automatically learning and extracting features. However, existing deep learning-based methods also have problems such as low algorithm performance and weak feature extraction ability. Summary of the Invention

[0004] To address the above problems, the present invention provides a deep clustering method for multi-feature fusion of asymmetric time-series data. It can not only receive the represented data for clustering, but also adapt to the characteristics of the original time-series data for clustering. Regarding the noise problem, while using a recurrence plot to transform the time-series data, the influence of noise is eliminated. The recurrence plot is constructed based on a sliding window. During this process, local noise is smoothed or discarded, and global features are retained. Regarding the complex features of time-series data, a recurrence plot plus a convolutional layer CNN and a long short-term memory network LSTM are respectively used to extract different features of the data for fusion, increasing the richness of features. An asymmetric variational autoencoder structure is adopted to increase the complexity of the encoder while reducing the complexity of the decoder, achieving the effect of enhancing the feature extraction ability.

[0005] To achieve the above object, the present invention discloses a deep clustering method for multi-feature fusion of asymmetric time-series data, including the following steps:

[0006] Step S1: Obtain time-series data, input the time-series data into the encoder LSTM to obtain the first feature, convert the time-series data into a recurrence plot RP, input the recurrence plot RP into the convolutional layer, transform the dimension of the value output by the convolutional layer to be the same as that of the first feature to obtain the second feature, and fuse the first feature and the second feature through a fusion layer;

[0007] Step S2: Sample the fusion layer to obtain the mean μ, variance σ, randomly sample ε from N(0,1) to obtain the embedded layer data z,

[0008] z = μ + σ * ε

[0009] Construct the loss function l of the embedding module r ,

[0010]

[0011] where x is the time-series data, is the output of the shallow decoder, is the reconstruction term loss function;

[0012] Step S3: Predict the probability that each sample belongs to each cluster to obtain the model prediction P, and construct the target variable Q based on the model prediction P;

[0013] Step S4: Define the empirical distribution F of the target variable Q, minimize the KL divergence between the empirical distribution F and the uniform distribution u, construct the loss function l of the regularization clustering layer c , and combine it with the loss function l of the embedding module r to obtain the final model loss function L;

[0014] Step S5: Train the clustering model based on the loss function L and the target variable Q to obtain a new target variable. Calculate the difference between the new target variable and the previous target variable. If it is less than the preset value or the number of iterations is greater than the maximum number of iterations, proceed to Step S6. If it is greater than the preset value, return to Step S3;

[0015] Step S6: Obtain the clustering result based on the new target variable:

[0016] CL = argmax i q ia

[0017] where argmax i q ia is to take a maximum value from each row of samples i, and a total of I values are obtained.

[0018] Further, Step S1 specifically includes:

[0019] The fusion layer e f is:

[0020] e f = e c + αe l

[0021] where e c is the output after dimensional transformation of the convolutional layer, e l is the output of the LSTM layer, and α is the weight.

[0022] Further, Step S3 specifically includes:

[0023] S31: Predict the probability P that each embedded layer data belongs to each cluster. All p ia constitute P:

[0024]

[0025] where z i is the embedded layer data of the i-th row of samples, A is the total number of cluster categories, a, a' are any cluster categories, N is the total number of Softmax function nodes, n is any Softmax function node, and θ n is the parameter of any node of the Softmax function;

[0026] S32: Construct the target variable Q. All q ia constitute Q:

[0027]

[0028] where I is the total number of sample rows, and i, i' are any row of samples.

[0029] Furthermore, step S4 specifically includes:

[0030] Construct the empirical distribution F of the target variable Q, and all f a constitute F:

[0031]

[0032] Minimize the KL divergence between the empirical distribution F and the uniform distribution u, and construct the loss function l of the regularization clustering layer c :

[0033] l c = KL(P‖Q)+KL(F‖u)

[0034] Combine the loss function l of the embedding module r , and obtain the final model loss function L:

[0035]

[0036] The beneficial effects of a multi-feature fusion asymmetric time-series data deep clustering method of the present invention are as follows: It can not only receive the represented data for clustering, but also adapt to the characteristics of the original time-series data for clustering. For the noise problem, while using the recursive graph to transform the time-series data, the influence of noise is eliminated. The recursive graph is constructed based on a sliding window. During this process, local noise is smoothed or discarded, and global features are retained. For the complex features of time-series data, the recursive graph plus the convolutional layer CNN and the long short-term memory network LSTM are respectively used to extract different features of the data for fusion, increasing the richness of features. The asymmetric variational autoencoder structure is adopted to increase the complexity of the encoder while reducing the complexity of the decoder, achieving the effect of enhancing the feature extraction ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0038] Figure 1 is a schematic diagram of the multi-feature fusion framework of the present invention.

[0039] Figure 2 is a schematic diagram of the clustering model structure of the present invention.

[0040] Figure 3 is a schematic diagram of the CNN structure of the present invention.

[0041] Figure 4 is a schematic diagram of the global recursive graph structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0043] The present invention proposes a multi-feature fusion asymmetric time series data deep clustering model FFADC (Featurefused asymmetric deep clustering, FFADC). As Figure 2 shown, it is a schematic diagram of the FFADC model structure. It can not only receive the represented data for clustering, but also adapt to the characteristics of the original time series data for clustering. Regarding the noise problem, FFADC uses a recursive graph to transform the time series data while eliminating the influence of noise. This is because the distribution of noise is discrete and local. The recursive graph is constructed based on a sliding window. During this process, local noise is smoothed or discarded, and global features are retained. And because the deep neural network has the characteristics of multiple abstraction levels and parallel computing, it has high adaptability and computing efficiency when processing high-dimensional and large-scale data. Regarding the complex features of time series data, a recursive graph plus a convolutional layer CNN (Convolutional Neural Network) and a long short-term memory network LSTM (LongShort-Term Memory) are respectively used to extract different features of the data for fusion, increasing the richness of features. And an asymmetric variational autoencoder structure is adopted to increase the complexity of the encoder while reducing the complexity of the decoder, achieving the effect of enhancing the feature extraction ability. In addition, by adding a clustering layer, the embedding and clustering jointly participate in the network training. Experiments show that FFADC can achieve better clustering results than other time series data clustering algorithms.

[0044] The representation of time series data can be applied to multiple aspects such as prediction, anomaly detection, and decision-making. Clustering also plays an important role in this. A good representation of time series data helps to improve the clustering effect. However, the representation is independent of clustering and only shows the structure and features of the data. The ability of the clustering algorithm to capture these features is also very important. Existing algorithms have certain limitations in time series data clustering. In order to adapt to different scenarios and obtain better clustering results, in-depth research has been carried out on the clustering algorithm for time series data. Regarding the characteristics of time series data and combining deep learning, the FFADC algorithm is proposed.

[0045] An embodiment of the present invention discloses a multi-feature fusion asymmetric time series data deep clustering method, including the following steps:

[0046] Step S1. To enable the features extracted by the model to contain richer information, this method uses different neural network layers to extract features from the time series data. The schematic diagram of the multi-feature fusion framework is as shown in Figure 1 Figure [not provided]. Two types of features are respectively extracted by the recurrence plot, the CNN structure, and the LSTM. First, the time series data is converted into a recurrence plot RP (Recurrence plot). The recurrence plot contains features such as the period and trend of the time series data and can eliminate local noise in the data. Subsequently, the recurrence plot is input into the convolutional layer to obtain the first part of the features. Then, the time series data is input into the LSTM to obtain the second part of the features.

[0047] Since the outputs of the convolutional layer and the LSTM layer have different dimensions, in order to fuse the two, it is necessary to flatten the output of the convolutional layer and transform it to the same dimension as the LSTM output through a fully connected layer. Let e c represent the output of the convolutional layer after dimensional transformation, and e l represent the output of the LSTM layer. The fusion layer e f is obtained by the formula e f = e c + αe l .

[0048] Among them, α is used to determine the weights of the two types of features during fusion. The current value is 1, that is, the two types of features have the same weight.

[0049] During the derivation of the regularization term, the method of variational inference is adopted, and it is necessary to optimize the Evidence Lower Bound (ELBO), as shown in the following formula:

[0050] ELBO(q) = logp(e f ) - KL(q(z)‖p(z|e f ))

[0051] Among them, p(e f ) represents the distribution sampled from the data e f of the fusion layer. logp(e f ) is a constant here. z represents the data of the embedding layer, that is, the extracted data features. q(z) represents the distribution sampled from the embedding layer data z. p(z|e f ) represents the distribution obtained from the data e fGenerate the distribution of the embedded layer data z. KL represents the Kullback-Leibler divergence, and all variables in the formula can be calculated. However, if directly sampling q(z), it will cause the neural network backpropagation to be unable to update. Therefore, the reparameterization trick is used. First, directly sample the mean μ and variance σ from the fusion layer, and randomly sample ε from N(0,1). At this time, the sampling of q(z) can be expressed by the following formula:

[0052] z = μ + σ * ε

[0053] Step S2: As described above, sample the fusion layer to obtain the mean μ, variance σ, and randomly sample ε from N(0,1) to obtain the embedded layer data z,

[0054] z = μ + σ * ε

[0055] Construct the loss function l of the embedding module r ,

[0056]

[0057] where x is the time series data, is the output of the shallow decoder based on the embedded layer z, is the loss function of the reconstruction term;

[0058] Step S3: Predict the probability that each sample belongs to each cluster to obtain the model prediction P, and construct the target variable Q based on the model prediction P;

[0059] Step S4: Define the empirical distribution F of the target variable Q, minimize the KL divergence between the empirical distribution F and the uniform distribution u, construct the loss function l of the regularization clustering layer c , and combine it with the loss function l of the embedding module r to obtain the final model loss function L;

[0060] Step S5: Train the clustering model based on the loss function L and the target variable Q to obtain the clustering model. Calculate the difference between the clustering center of this clustering model and the previous clustering center. If the difference is less than the preset value or the number of iterations is greater than the maximum number of iterations, then perform Step S6. If it is greater than the preset value and the number of iterations is less than the maximum number of iterations, then return to Step S3 and use the gradient descent method to recalculate P and Q for the next round;

[0061] Step S6: Obtain the clustering result based on the new target variable:

[0062] CL = argmax i q ia

[0063] where, argmax i qia For each row of samples \(i\), take the largest \(q\) ia value, and the \(a\) corresponding to the largest \(q\) ia value is the cluster category to which the corresponding sample belongs.

[0064] To further optimize the above technical solution, step S1 specifically includes:

[0065] Fusion layer \(e\) f is:

[0066] \(e\) f \(=e\) c \(+\alpha e\) l

[0067] where \(e\) c is the output after dimensional transformation of the convolutional layer, \(e\) l is the output of the LSTM layer, and \(\alpha\) is the weight.

[0068] To further optimize the above technical solution, step S3 specifically includes:

[0069] S31. In order to obtain stable results in a complex feature space, the present invention constructs a new regularization clustering module, which uses multiple softmax nodes and retains the regularization term (to prevent the vast majority of samples from being assigned to one or a few clusters). In step S3, first predict the probability \(P\) that each data in the embedding layer belongs to each cluster. All \(p\) ia constitute \(P\):

[0070]

[0071] where \(z\) i is the data of the embedding layer of the \(i\)-th row of samples, \(A\) is the total number of cluster categories, \(a, a'\) are the numbers of any cluster categories, \(N\) is the total number of softmax function nodes, \(n\) is any softmax function node, and \(\theta\) n is the parameter of any node of the softmax function;

[0072] The total number of cluster categories \(A\), the total number of softmax function nodes \(N\), and the total number of sample rows \(I\) are all determined by the time-series data \(X\);

[0073] S32. To define the clustering loss function, it is necessary to use the auxiliary variable \(Q\) (hereinafter referred to as the target variable) to iteratively optimize the model prediction. All \(q\) ia constitute \(Q\):

[0074]

[0075] where \(I\) is the total number of sample rows, and \(i, i'\) are any row of samples.

[0076] To further optimize the above technical solution, step S4 specifically includes:

[0077] Construct the empirical distribution F of the target variable Q, and all f a constitute F, where f a represents the assignment probability of any one cluster and is calculated by the following formula:

[0078]

[0079] Minimize the KL divergence between the empirical distribution F and the uniform distribution u, and construct the loss function l of the regularization clustering layer c :

[0080] l c = KL(P‖Q)+KL(F‖u)

[0081] Currently, most deep clustering algorithms perform embedding and clustering in two independent steps. The disadvantage of this approach is that the model will disrupt the spatial distribution of data during training clustering, resulting in a decrease in clustering accuracy. In the present invention, feature extraction and clustering are jointly trained, enabling the model to learn a feature distribution more suitable for clustering. Therefore, in the present invention, based on l c in combination with the loss function l r of the embedding module, the final model loss function L is obtained:

[0082]

[0083] The time series data acquisition methods in different scenarios will vary according to specific application requirements and the characteristics of data sources. For the time series data acquisition requirements in different scenarios, there are the following time series data acquisition methods.

[0084] 1. Sensor data acquisition: In the Internet of Things and sensor networks, time series data of various physical quantities are collected by deploying sensor devices, such as temperature, humidity, pressure, light, etc. The sensors can directly obtain real-time data in the environment and store or transmit it to the central server for further processing and analysis.

[0085] 2. Log data acquisition: In the fields of computer systems, networks, and servers, etc., log files are generated by recording system operations and events for subsequent fault diagnosis, performance optimization, etc. analysis. The log data contains timestamps and event information and can be used to construct time series data for analysis.

[0086] 3. Instrument measurement data acquisition: In scientific experiments and engineering tests, various instrument devices are used for measurement, such as electrocardiograms, electroencephalograms, sound waveforms, etc. The instruments will acquire data at a certain sampling frequency and generate time series data for further analysis and processing.

[0087] 4. Financial Market Data Collection: In the financial field, financial market data such as stock prices, foreign exchange rates, bond yields, etc. are collected through channels such as financial exchanges or financial data providers. These data are provided in the form of time series and are used for financial analysis, investment decision-making, etc.

[0088] 5. Social Media Data Collection: On social media platforms, data such as messages, comments, and forwards posted by users are collected through API interfaces or web crawler technologies. These data contain time information and can be used to construct time series data of user behavior for tasks such as social network analysis and sentiment analysis.

[0089] Convolutional Neural Networks (CNNs) are a type of neural network model used in the field of data processing for data with grid structures such as images, speech, and videos. Convolutional neural networks typically consist of a convolutional layer, a pooling layer, and a fully connected layer, etc. Among them, the convolutional layer is the core of the convolution operation, which can extract various features from images, such as edges, textures, shapes, etc. The pooling layer can reduce the size of the feature map, thereby reducing the computational amount and the risk of overfitting. Finally, the fully connected layer is used to map the features to the output categories to complete classification or prediction tasks. Compared with traditional fully connected neural networks, CNNs use techniques such as convolutional layers, pooling layers, and non-linear activation functions to improve the recognition accuracy and generalization ability of the model. Its schematic diagram is as Figure 3 shown;

[0090] The operation of taking the inner product (multiplying element by element and then summing) of an image and a filter matrix filter is the so-called convolution operation. In CNNs, the convolution filter filter performs a convolution operation on local input data, then moves the sliding window and continues to calculate until all data are calculated and processed.

[0091] The pooling operation takes the overall statistical features of an adjacent area at a certain position of the input matrix as the output of that position. There are mainly average pooling and max pooling, etc. Simply put, pooling is to specify a value in this area to represent the entire area.

[0092] CNNs usually involve the following hyperparameters:

[0093] (1) Filter Size: Defines the size of the convolution kernel, usually an odd number, such as 3x3, 5x5, etc.

[0094] (2) Stride: Defines the stride of the convolution kernel movement.

[0095] (3) Zero-padding: Add extra meaningless pixels around the edges of the input to reduce the loss of edge information.

[0096] (4) Number of convolutional kernels: The number of convolutional kernels to be trained, and each convolutional kernel performs an inner product operation separately.

[0097] (5) Activation function: The activation function for each convolutional layer, such as ReLU, Sigmoid, etc.

[0098] (6) Pooling: The type and size of the pooling layer used for downsampling, such as max pooling, average pooling, etc.

[0099] (7) Batch normalization: Normalize the output of each layer to ensure the stability of the entire network.

[0100] Due to its unique structure, CNN has two characteristics: local connection and weight sharing. Local connection: The convolutional neural network uses convolutional kernels to extract local features of the input. At the same time, different convolutional kernels are used for different features, which can effectively reduce the number of parameters.

[0101] Weight sharing: Neurons in the same layer use the same weights, reducing the number of neurons and the risk of overfitting.

[0102] CNN has a wide range of applications in the fields of image, speech, and video recognition, such as face recognition, license plate recognition, natural language processing, etc. The advantages of CNN can automatically learn and extract features without manual feature design, reducing the human and time costs. At the same time, CNN has the ability of parallel computing, which can efficiently handle the training and prediction tasks of large-scale data sets, thus improving the response speed of the system.

[0103] The recurrence plot is a tool for visualizing nonlinear dynamical systems. It can explain the internal structure of time series, give prior knowledge about similarity, information content, and predictability, and is an important method for analyzing the periodicity, trend, chaos, and non-stationarity of time series. The recurrence plot represents the recurrence of the trajectory of the dynamical system in the phase space. The steps for converting time series data into a recurrence plot are as follows:

[0104] (1) Given time series data Determine the embedding dimension m and the delay τ, and reconstruct it. The reconstructed one Can be represented by the following formula, where i ∈ [1, n - (m - 1)τ]:

[0105]

[0106] (2) The recurrence graph is obtained through the following formula:

[0107]

[0108] where θ represents the Heaviside function, ε is the recurrence threshold, representing the maximum acceptable value at which two trajectories can be regarded as periodic. At this time, the value of R ij is 0 or 1. If the Heaviside function θ is removed, as shown in the following formula, it is called the global recurrence graph. The value of R ij is between 0 and 1. The global recurrence graph is adopted in this model:

[0109]

[0110] After the time series data is converted into the global recurrence graph, it is as Figure 4 shown.

[0111] VAE converts the data into a distribution rather than a specific value. Its loss is jointly composed of a reconstruction term and a regularization term. The reconstruction term is responsible for minimizing the difference between the input and output. The regularization term makes the distribution returned by the encoder close to the standard normal distribution and standardizes the organization of the latent space.

[0112] In the model, the CNN part of the encoder consists of three convolutional layers. The numbers of convolutional kernels are (32, 64, 128) respectively. The sizes of the convolutional kernels in each convolutional layer are the same, which are (5, 5, 3) respectively, and the stride is 2 for each layer. In order to avoid the problem of gradient disappearance and make the model non-linear, the ReLU activation function is adopted for each layer. The LSTM part adopts three stacked LSTMs, and the number of hidden units is set to 32. The length of the embedding layer is set to be the same as the number of categories. The shallow decoder is a structure in which a layer of LSTM is connected to a fully connected layer. The number of LSTM hidden units is set to 16, and the fully connected layer transforms the data dimension to be the same as the input data. The number of Softmax function nodes is 6, and the ADAM optimizer is used to iteratively train the model. In order to balance the stability and efficiency during model training, the mini-batch gradient descent method is adopted. The batch size is set to 32, the maximum number of iterations is 3000, and the target variable Q is updated every 140 iterations. It is judged whether the difference from the target variable before this iteration is less than the preset threshold. One of the conditions for the end of model training is to reach the maximum number of iterations, and the other is that the clustering center obtained in this iteration is less different from the previous one than a certain threshold. In this experiment, the threshold is set to 0.001.

[0113] Specifically, the gradient descent method is as follows:

[0114] 1. Objective function (loss function): First, we have an objective function to be minimized, usually denoted as J(θ), where θ is the parameter vector we want to adjust. Our goal is to find θ that minimizes the objective function J(θ).

[0115] 2. Initial parameters: Start with an initial parameter θ, usually initialized with random values or according to some heuristic method.

[0116] 3. Calculate the gradient: Calculate the gradient of the objective function J(θ) with respect to the parameter θ. The gradient represents the rate of change of the function at the current parameter value, and the gradient is usually denoted by where denotes the gradient operator.

[0117] 4. Update the parameters: Use the gradient information to update the parameter θ to minimize the objective function. The update rule is usually:

[0118]

[0119] where β is the learning rate, a hyperparameter that controls the step size. is the gradient calculated at the current parameter θ old and θ new and θ old correspond to θ in the above text n . The learning rate determines the step size of the parameter update in each iteration.

[0120] 5. Repeat the iteration: Repeat steps 3 and 4 until the stopping condition is met, such as reaching the maximum number of iterations, the objective function converges, etc.

[0121] The pseudo-code for the clustering step is shown in Algorithm 1:

[0122]

[0123] The textual description is as follows:

[0124] (1) Input the time series data into the embedding module of FFADC for pre-training to obtain the embedded layer data z.

[0125] (2) Apply the K-means algorithm to the embedded layer data z to obtain the initial clustering assignment, i.e., the probability P that sample i belongs to class a, and calculate the target variable Q using formula (6).

[0126] (3) Start training the FFADC model and iteratively optimize the model parameters.

[0127] (4) Calculate the difference between the current clustering assignment and the previous iteration's clustering assignment in each iteration. When the difference is less than the set threshold or the maximum number of iterations is reached, stop training to obtain the clustering result.

[0128] The proposed deep clustering model FFADC will be compared with other deep clustering models. The evaluation metrics are the silhouette coefficient and the CH index. The larger the value, the better the clustering effect. The deep clustering models for comparison are the existing deep clustering algorithms for time series data, CCAE, BRAC, and TS-SANC. At the same time, the data represented by FSR is used as the dataset and input into FFADC to obtain further clustering results.

[0129] The following table shows the comparison of the silhouette coefficient and the CH index of the clustering effects of each algorithm on different datasets:

[0130] Comparison of the silhouette coefficients of each algorithm on different datasets

[0131]

[0132] Comparison of the CH indices of each algorithm on different datasets

[0133]

[0134] It can be seen from the above table that on most datasets, FFADC has achieved the highest silhouette coefficient and CH index scores. Although other deep clustering models have demonstrated the advantages of deep neural networks in feature extraction and dimensionality reduction, their performance is still not ideal enough because they are not fully designed for the characteristics of time series data. For the FSR dataset, after clustering by inputting it into FFADC, its silhouette coefficient and CH index have also been improved to a certain extent.

[0135] The schematic diagram of the clustering results of FFADC on each dataset after visualization shows that each sample is clearly classified, and only a very small number of sample points are classified into the wrong clusters.

[0136] The present invention proposes an asymmetric deep clustering method for multi-feature fusion of time series data, which can not only receive the represented data for clustering, but also adapt to the characteristics of the original time series data for clustering. To address the noise problem, FFADC uses a recursive graph to transform the time series data while eliminating the influence of noise. This is because the distribution of noise is discrete and local, and the recursive graph is constructed based on a sliding window. During this process, local noise is smoothed or discarded, and global features are retained;

[0137] Deep neural networks have multiple levels of abstraction and parallel computing, and have high adaptability and computing efficiency in processing high-dimensional and large-scale data. To address the complex features of time series data, a recursive graph plus a convolutional layer CNN (Convolutional Neural Network) and a long short-term memory network LSTM (Long Short-Term Memory) are respectively used to extract different features of the data for fusion, increasing the richness of features;

[0138] An asymmetric variational autoencoder structure is adopted to increase the complexity of the encoder while reducing the complexity of the decoder, achieving the effect of enhancing the feature extraction ability. In addition, by adding a clustering layer, the embedding and clustering jointly participate in the network training, and the pattern recognition ability of the deep neural network is used to further optimize the clustering results.

[0139] Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

Claims

1. An asymmetric time series data deep clustering method with multi-feature fusion, characterized in that Including the following steps: Step S1: Obtain time series data, input the time series data into the encoder LSTM to obtain a first feature, convert the time series data into a recurrence plot RP, input the recurrence plot RP into a convolutional layer, transform the dimension of the value output by the convolutional layer to the same dimension as the first feature to obtain a second feature, and fuse the first feature and the second feature through a fusion layer; Step S2: Sample the fusion layer to obtain a mean μ, a variance σ, randomly sample ε from N(0, 1) to obtain embedded layer data z, z = μ + σ * ε Construct the loss function l of the embedding module r , where x is the time series data, is the output of the shallow decoder, is the reconstruction term loss function; Step S3: Predict the probability that each sample belongs to each cluster to obtain a model prediction P, and construct a target variable Q based on the model prediction P; Step S4: Define the empirical distribution F of the target variable Q, minimize the KL divergence between the empirical distribution F and the uniform distribution u, and construct the loss function l of the regularization clustering layer c , and combine it with the loss function l r of the embedding module to obtain the final model loss function L; Step S5: Train the clustering model based on the loss function L and the target variable Q to obtain a new target variable, calculate the difference between the new target variable and the previous target variable. If it is less than a preset value or the number of iterations is greater than the maximum number of iterations, then perform step S6. If it is greater than the preset value, then return to step S3; Step S6: Obtain a clustering result based on the new target variable: CL = argmax i q ia Among them, argmax i q ia is to take a maximum value from each row of sample i, and a total of I values are obtained.

2. A deep clustering method for asymmetric time-series data with multi-feature fusion according to claim 1, characterized in that, The specific content of step S1 includes: Fusion layer e f is as follows: e f = e c + αe l Among them, e c is the output after the convolutional layer undergoes dimensional transformation, and e l is the output of the LSTM layer, where α is the weight.

3. A method for deep clustering of asymmetric time series data with multi-feature fusion according to claim 1, characterized in that, The specific content of step S3 includes: S31. Predict the probability P that the data of each embedding layer belongs to each cluster. All p ia constitute P: where z i is the embedded layer data of the i-th row of samples, A is the total number of cluster categories, a and a' are any cluster categories, N is the total number of Softmax function nodes, n is any Softmax function node, and θ n is the parameter of any node of the Softmax function; S32. Construct the target variable Q, all q ia constitute Q: Where I is the total number of sample rows, and i, i′ are any row of samples.

4. A method for deep clustering of asymmetric time series data with multi-feature fusion according to claim 3, characterized in that The specific content of step S4 includes: Construct the empirical distribution F of the target variable Q, all f a Constitute F: Minimize the KL divergence between the empirical distribution F and the uniform distribution u to construct the loss function l of the regularization clustering layer c : l c = KL(P||Q) + KL(F||u) Loss function \(l\) combined with the embedding module r , resulting in the final model loss function \(L\):

Citation Information

Cited By

  • Multivariable time series data identification method, equipment and medium

    CN120873817A