Time sequence clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion

By constructing a time series clustering method with multi-granularity data enhancement and multi-domain reliable class diffusion, generating rich time domain and frequency domain enhanced data, designing a multi-domain pattern extractor and implementing intra-class and inter-class diffusion mechanisms, the clustering accuracy problem of traditional time series clustering methods under high-dimensional nonlinear characteristics is solved, achieving more efficient wind condition forecasting and stable operation of wind power generation systems.

CN120744554APending Publication Date: 2025-10-03DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510977370.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional time series clustering methods have insufficient clustering accuracy under high-dimensional, nonlinear and time-varying features. They are difficult to effectively distinguish semantically similar samples and have low utilization of data enhancement information, resulting in limited model robustness. They also ignore the semantic correlation between samples within a class, resulting in waste of semantic information.

Method used

A data augmentation strategy based on mixed multi-granularity perturbations is constructed to generate enhanced data in the time and frequency domains. A clustering-guided multi-domain pattern extractor is designed to extract the clustering probability distribution of original data and enhanced data in the time and frequency domains. Reliable clustering of time series samples is achieved through a multi-domain reliable class diffusion strategy, including intra-class and inter-class double diffusion mechanisms to enhance clustering stability and reliability.

Benefits of technology

The accuracy and stability of time series clustering have been improved, which can better identify wind operating modes and improve the operating efficiency and reliability of wind power generation systems in complex mountainous wind farm environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744554A_ABST
    Figure CN120744554A_ABST
Patent Text Reader

Abstract

The invention provides a time sequence clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion, and belongs to the technical field of time sequence data analysis, modeling is performed on an unsupervised meteorological time sequence data clustering algorithm based on a depth comparison architecture of cross-view reliable class diffusion, and the method comprises three parts, and constructing a data enhancement strategy based on mixed multi-granularity disturbance, designing a multi-domain mode extractor based on clustering guidance, and constructing a multi-domain reliable class diffusion strategy. According to the method, intra-class compactness and inter-class separability are enhanced based on a multi-domain isomorphic space class diffusion strategy, and the time sequence clustering capability of the model is improved; a multi-domain mode extractor based on clustering guidance is designed, clustering probability distribution of original data and enhanced data in a time domain and a frequency domain is extracted, and it is ensured that data of different domains have comparability in the same space; and a multi-domain reliable class diffusion strategy is constructed, similar samples in an isomorphic space are promoted to be aggregated, heterogeneous samples are promoted to be separated, and the clustering stability and reliability are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of time series data analysis, and relates to a time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion. Background Art

[0002] In the complex and ever-changing wind farm environment of mountainous areas, accurate time-series wind forecasting is crucial for ensuring the efficient operation of wind power generation systems. Transient fluctuations in wind speed and direction not only directly impact aerodynamic load distribution and turbine control strategies, but also determine the accuracy limits of ultra-short-term power forecasts. However, due to insufficient resolution, traditional technologies struggle to capture these subtle and rapid changes, resulting in inaccurate understanding and prediction of wind farms. This inapplicability not only reduces wind turbine performance and power generation efficiency, but can also increase mechanical loads and wear on wind turbines due to miscalculation of wind conditions, shortening equipment life and, in extreme cases, even causing equipment damage.

[0003] To solve this problem, it is crucial to deeply mine the effective information in high-resolution wind measurement data. Time series clustering technology shows its unique advantages at this time. It can deeply mine the temporal and spatial correlation patterns implicit in high-resolution wind measurement data and achieve hierarchical decoupling of the evolution laws of multi-dimensional wind conditions. Through cluster analysis, wind condition data with similar change trends are classified into one category, so that different wind farm operating modes can be clearly identified. Based on this, an intelligent analysis framework that can adaptively identify wind farm operating modes can be constructed, providing key feature support for the construction of scenario-driven dynamic prediction models, thereby effectively improving the accuracy of wind condition time series predictions and better meeting the operation requirements of wind power generation systems in complex mountainous wind farm environments.

[0004] In recent years, unsupervised time series data clustering methods based on deep learning have made great progress, which mainly include two categories: non-contrastive and contrastive. Non-contrastive methods drive deep networks to perform feature extraction and cluster assignment by designing predefined optimization objectives such as reconstruction loss functions and time series invariance constraints. Although this type of method can capture basic time series features, it has two significant defects: first, the feature discriminability is insufficient, and it is difficult to effectively distinguish semantically similar samples; second, the utilization rate of data enhancement information is low, resulting in limited model robustness. Contrastive deep clustering methods demonstrate advantages in improving feature discriminability by constructing positive and negative sample pairs to implement instance discriminant learning. However, existing contrastive methods still have the following key technical defects:

[0005] First, the encoding networks designed for traditional contrastive clustering modules typically employ a two-branch feature output: a high-dimensional feature representation that retains fine-grained instance information for instance-level contrastive learning, strengthening the discriminability of the samples themselves by measuring the similarities between individual samples; and a low-dimensional feature representation that, after dimensionality compression, focuses on the clustering properties of the samples for cluster-level contrastive learning, optimizing the discriminability of cluster boundaries by measuring the distribution differences between different samples (such as cluster center distances). While these designs can capture both inter-sample differences and inter-cluster distinguishability, the lack of a cross-space hierarchical feature interaction mechanism makes it difficult to establish a bidirectional knowledge flow path between local sample features and the global clustering structure.

[0006] Secondly, traditional instance-level contrastive learning uses a prototype-centric rigid negative sampling mechanism, which uniformly treats non-prototype samples as negative samples and rejects them. While this strategy strengthens prototype discrimination, it ignores the semantic relevance between samples within a class, resulting in an inability to effectively utilize the potential similarities between non-prototype samples of the same class, resulting in a waste of semantic information.

[0007] In summary, the present invention proposes a time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion, constructs a data enhancement strategy based on hybrid multi-granularity perturbation to mine enhanced data in the time domain and frequency domain, designs a clustering-oriented multi-domain pattern extractor to extract the clustering probability distribution of original data and enhanced data in the time domain and frequency domain, and finally constructs a multi-domain reliable class diffusion strategy based on the intra-class and inter-class double diffusion mechanism to achieve reliable clustering of time series samples. Summary of the Invention

[0008] The present invention proposes a time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion, aiming to effectively solve the clustering accuracy problem of traditional meteorological data clustering methods under high-dimensional, nonlinear and time-varying features. The present invention constructs a data enhancement strategy based on mixed multi-granularity perturbations to obtain enhanced data in the time domain and frequency domain; designs a multi-domain pattern extractor based on clustering guidance to extract the clustering probability distribution of original data and enhanced data in the time domain and frequency domain; constructs a multi-domain reliable class diffusion strategy to perform a reliable class diffusion process in the time domain, frequency domain and cross-domain, so as to make similar samples closer and heterogeneous samples more distant, thereby improving the clustering effect. After training is completed, the original time series data is input into the model to obtain the final meteorological data category.

[0009] In order to achieve the above object, the technical solution adopted by the present invention is:

[0010] Step 1: Construct a data enhancement strategy based on hybrid multi-granularity perturbations to generate enhanced data in the time and frequency domains. In the time domain, a multi-granularity time domain data enhancement strategy is constructed by combining jitter, scaling, and permutation operations to generate time domain enhanced data that preserves the local trends and global structure of the time series data. In the frequency domain, the original time domain data is first converted into original frequency domain data through Fourier transform. Then, a multi-granularity frequency domain data enhancement strategy is constructed by combining frequency addition and removal operations to generate frequency domain enhanced data that covers the diverse variation patterns of frequency features.

[0011] Step 2: Design a cluster-guided multi-domain pattern extractor to extract the cluster probability distribution of the original and enhanced data in the time and frequency domains, ensuring that data from different domains are comparable in the same space. The time-domain pattern extractor uses a bidirectional long short-term memory network and a clustering block cascade architecture, while the frequency-domain pattern extractor uses a three-layer convolutional block and a clustering block cascade architecture.

[0012] Step 3: Build a multi-domain reliable class diffusion strategy. By designing a dual diffusion mechanism for intra-class and inter-class data, we construct a triple-class diffusion loss function for time, frequency, and cross-domain data to promote reliable clustering of time series data. The intra-class diffusion mechanism enhances clustering stability by strengthening the similarity between samples and their nearest reliable neighbors in the latent space, while the inter-class diffusion mechanism maximizes the differences between cluster representations to achieve reliable separation between clusters.

[0013] The beneficial effects of the present invention are:

[0014] The present invention constructs a data enhancement strategy based on mixed multi-granularity perturbations to generate rich enhanced data in the time and frequency domains to support the clustering learning process of high-dimensional nonlinear meteorological data. A cluster-oriented multi-domain pattern extractor is designed to extract the clustering probability distribution of original data and enhanced data in the time and frequency domains, ensuring that data from different domains are comparable in the same space. A multi-domain reliable class diffusion strategy is constructed to promote the aggregation of similar samples and the separation of heterogeneous samples in homogeneous spaces, thereby enhancing clustering stability and reliability. Finally, through unsupervised time series clustering experiments, it is verified that the present invention can achieve good time series clustering effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a framework diagram of the time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion.

[0016] Figure 2 This is the basic flow chart of the temporal clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion. DETAILED DESCRIPTION

[0017] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0018] like Figure 1 and Figure 2 As shown, the present invention is a time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion, which includes three parts: constructing a data enhancement strategy based on hybrid multi-granularity perturbations, constructing a multi-domain pattern extractor based on clustering guidance, and constructing a multi-domain reliable class diffusion strategy. First, a data enhancement strategy based on hybrid multi-granularity perturbations is constructed to generate enhanced data in the time domain and frequency domain. Then, a multi-domain pattern extractor based on clustering guidance is designed to extract the clustering probability distribution of original data and enhanced data in the time domain and frequency domain to ensure that data in different domains are comparable in the same space. Finally, a multi-domain reliable class diffusion strategy is constructed. By designing a double diffusion mechanism within and between classes, a triple class diffusion loss function in the time domain, frequency domain and cross-domain is constructed to promote the reliable clustering of time series data. The specific implementation of each step is as follows:

[0019] Step 1: Construct a data enhancement strategy based on mixed multi-granularity perturbations;

[0020] The data enhancement strategy based on hybrid multi-granularity perturbations consists of a multi-granularity time domain data enhancement strategy and a multi-granularity frequency domain data enhancement strategy. In the time domain dimension, a multi-granularity time domain data enhancement strategy is constructed by combining jitter, scaling, and permutation operations, and finally generates time domain enhanced data to preserve the local trend and global structure of the time series data; in the frequency domain dimension, the time domain raw data is first converted into frequency domain raw data through Fourier transform, and then the frequency addition and removal operations are combined to construct a multi-granularity frequency domain data enhancement strategy, and finally generates frequency domain enhanced data to cover the diverse variation patterns of frequency features. The details are as follows:

[0021] Step 1.1, construct a multi-granularity temporal data enhancement strategy;

[0022] Given time domain raw data , represents the number of samples, Indicates the number of clusters (categories). express No. Time domain original samples, Indicates the sequence length of the sample. , by constructing a multi-granularity time domain data enhancement strategy to enrich its time domain information.

[0023] The multi-granularity time domain data enhancement strategy consists of jitter, scaling and permutation operations. The jitter operation adds random noise to the time series data. The scaling operation Performs element-wise multiplication of and a given scaling factor. The permutation operation first uses a split operation to The time dimension is divided into The multi-granularity temporal data augmentation strategy is described as follows:

[0024] (1)

[0025] in, express No. Time domain original samples; express Enhanced time domain enhanced samples; Indicates a dither operation; Represents a zoom operation; Represents a permutation operation; Represents random noise, which has a mean of 0 and a variance of Normal distribution ; represents the scaling factor, The mean is 2 and the variance is Normal distribution; Represents a split operation; express The number of segments divided; L represents The sequence length of Indicates random shuffling operation; Represents a new time series; Represents the time step index obtained after the permutation operation; Represents an element-wise multiplication operation.

[0026] Step 1.2, construct a multi-granularity frequency domain data enhancement strategy;

[0027] For time domain raw data Each row of time domain original samples adopts the multi-granularity time domain data enhancement strategy of step 1.1 and merges them to obtain time domain enhanced data However, focusing only on time domain information may overlook key features. Therefore, the present invention first analyzes the time domain raw data. Fast Fourier Transform Get the original frequency domain data Next, for The Original frequency domain samples , by constructing a multi-granularity frequency domain data enhancement strategy to generate frequency domain enhanced samples .

[0028] The multi-granularity frequency domain data enhancement strategy consists of a local frequency removal operation and a local frequency addition operation. The local frequency removal operation will The mask of the local frequency domain removal operation is multiplied element by element. The local frequency addition operation first generates a random perturbation, and then the random perturbation is multiplied element by element with the mask of the local frequency addition operation. The specific process of the multi-granularity frequency domain data enhancement strategy is as follows:

[0029] (2)

[0030] in, express The Original frequency domain samples; express Enhanced frequency domain enhanced samples; represents the local frequency removal operation; The mask representing the local frequency domain removal operation, The probability of success is Bernoulli distribution ,in represents the probability of retaining the frequency component; represents a local frequency addition operation; Represents a random perturbation that follows the interval (0,0.1⋅max( )) uniform distribution on; The mask representing the local frequency addition operation, The probability of success is Bernoulli distribution ; Represents a hyperparameter. In this invention, the hyperparameter The range is 0.1~0.5 (the value in this embodiment is 0.2).

[0031] Step 2: Design a clustering-guided multi-domain pattern extractor;

[0032] For time domain raw data and frequency domain raw data By respectively adopting the time domain data enhancement strategy of step 1.1 and the frequency domain data enhancement strategy of step 1.2, the time domain enhanced data can be obtained. and frequency domain enhanced data In order to achieve 、 、 and Cluster probability distribution extraction in homogeneous space, the present invention is based on step 1 、 、 and , a cluster-oriented multi-domain pattern extractor was constructed to extract the cluster probability distribution of these data. The extractor consists of a time domain pattern extractor and a frequency domain pattern extractor, which generate the original and enhanced probability distributions in the time domain and frequency domain respectively. The time domain pattern extractor uses a bidirectional long short-term memory network and a clustering block cascade architecture to generate the original probability distribution and enhanced probability distribution in the time domain; the frequency domain pattern extractor uses a three-layer convolution block and a clustering block cascade architecture to extract the original probability distribution and enhanced probability distribution in the frequency domain. The details are as follows:

[0033] Step 2.1, construct a time domain pattern extractor;

[0034] The time domain pattern extractor Adopting the architecture of bidirectional long short-term memory network and clustering block cascade, yes The bidirectional LSTM block consists of a forward and backward LSTM network in parallel, relying on a bidirectional propagation mechanism to capture the forward and backward dependencies of sequence data. Specifically, the forward LSTM network processes the input forward from the beginning to the end of the sequence, while the backward LSTM network processes the input in reverse order from the end to the beginning. The time-step hidden states of the two are concatenated to form a feature representation that integrates bidirectional contextual information. The clustering block consists of two fully connected layers and a soft allocation function.

[0035] The time domain raw data and temporal enhancement data Input to the time domain pattern extractor , we can get the original probability distribution in the time domain and the enhanced probability distribution in the time domain. The time domain pattern extraction process is shown in formula (3):

[0036] (3)

[0037] in, and represent the original probability distribution in the time domain and the enhanced probability distribution in the time domain respectively.

[0038] Step 2.2, construct a frequency domain pattern extractor;

[0039] The frequency domain pattern extractor Adopting the architecture of three-layer convolutional blocks and clustering blocks cascaded, yes learnable parameters. Within each convolutional block, the output of the convolutional layer undergoes batch normalization, rectified linear unit activation, and one-dimensional max pooling. The first convolutional block implements an additional regularization layer after the pooling layer to randomly discard pooled features, further improving model generalization. The clustering block consists of two fully connected layers and a soft allocation function.

[0040] The original frequency domain data and frequency domain enhanced data Input to the frequency domain pattern extractor , we can get the original probability distribution in the frequency domain and the enhanced probability distribution in the frequency domain. The frequency domain pattern extraction process is shown in formula (4):

[0041] (4)

[0042] in, and represent the original probability distribution in the frequency domain and the enhanced probability distribution in the frequency domain respectively.

[0043] Step 3: Build a multi-domain reliable diffusion strategy

[0044] For time domain raw data With time domain enhanced data Using the time domain pattern extractor in step 2.1, we can get the original probability distribution in the time domain Time-domain enhanced probability distribution . For the original frequency domain data With frequency domain enhanced data Using the frequency domain pattern extractor in step 2.2, we can get the original probability distribution in the frequency domain Frequency domain enhanced probability distribution In order to derive reliable time series clustering, based on the above 、 、 and , the present invention constructs a multi-domain reliable class diffusion strategy, which realizes reliable clustering of multi-domain data by designing intra-class diffusion mechanism and inter-class diffusion mechanism. First, a reliable intra-class diffusion mechanism is designed for the time domain, frequency domain and cross-domain. By constructing a sample neighborhood set, generating a semantic indicator matrix and designing an intra-class diffusion loss, the reliable similarity between the sample and its potential space neighbors is enhanced and the clustering stability is improved. Then, a reliable inter-class diffusion mechanism is designed for the time domain, frequency domain and cross-domain. By introducing an inter-class diffusion loss with the goal of maximizing partition entropy, reliable separation between clusters is achieved. Finally, the intra-class diffusion loss and inter-class diffusion loss of the time domain, frequency domain and cross-domain are integrated to construct a triple class diffusion loss function to achieve stable and reliable clustering. The details are as follows:

[0045] Step 3.1, design a reliable intra-class diffusion mechanism;

[0046] To simplify the description, ( and represent time domain and frequency domain respectively), define and Represents domain The original probability distribution and enhanced probability distribution of .

[0047] First construct and The similarity matrix between the two , establish cross-view semantic associations between samples through the similarity matrix. Then for No. Row vector , according to the similarity matrix exist Select the one with the greatest similarity Instances construct their sample neighborhood set At the same time, in order to alleviate the false positive problem caused by the introduction of domain sets, the present invention introduces the semantic indicator matrix Among them, the semantic indicator matrix Middle Rank Column element value As shown in formula (5):

[0048] (5)

[0049] in, represents the maximum assignment probability index generating function; express No. row vector, express No. Row vector.

[0050] Combining the neighborhood set and the semantic indicator matrix, we can further construct a cross-view neighborhood distillation matrix , which provides a reliable neighborhood relationship for inter-class diffusion. Among them, the cross-view neighborhood distillation matrix Middle Rank Column element value As shown in formula (6):

[0051] (6)

[0052] In formula (6), Samples within a neighborhood with consistent cluster predictions should be close together to suppress false positive sample interference in latent structure learning. Conversely, samples that do not meet the above conditions should remain relatively static to mitigate false negative sample interference in mining structural differences between samples. Intra-class reliability and compactness are achieved by measuring the similarity of samples within a neighborhood with consistent cluster predictions.

[0053] Finally, the intra-class diffusion loss of the design domain v is ( , )for:

[0054] (7)

[0055] in, ( , )express and Intra-class diffusion loss; is the row normalization operator, is the number of neighbors after distillation, Represents element-by-element multiplication operation, exp represents the exponential function with the natural constant e as the base, express The first Rank Elements of the column, express No. Row vector.

[0056] The formula (7) Replace with and After that, we can get the intra-class diffusion loss in the time domain ( , ) and frequency domain intra-class diffusion loss ( , ). In addition, in order to ensure the consistency of intra-class diffusion in time domain and frequency domain, the present invention also calculates the cross-domain intra-class diffusion loss. In formula (7), the input variable and Replace with and We can get the cross-domain intra-class diffusion loss ( , ).

[0057] Step 3.2, design a reliable inter-class diffusion mechanism;

[0058] The reliable intra-class diffusion mechanism design in step 3.1 is to enhance the intra-class compactness of time series clustering. In addition, in order to enhance the inter-class separation of time series clustering, the present invention then designs a reliable inter-class diffusion mechanism. Specifically, by The original probability distribution and enhanced probability distribution Imposing orthogonal constraints forces the semantic representation to approach the geometric simplex of the semantic space, thereby enhancing the ability to capture boundaries.

[0059] Will and Each column of is interpreted as a specific cluster representation, and then the clusters are reliably separated by maximizing the difference between the cluster representations. This introduces the inter-class diffusion loss with the goal of maximizing the partition entropy:

[0060] (8)

[0061] in, express and inter-class diffusion loss; express The mean probability that all instances in belong to the jth cluster; express The mean probability that all instances in belong to the jth cluster; represents the identity matrix; represents information entropy loss; express No. The clusters represent express No. The clusters represent express and The similarities between the two.

[0062] The formula (8) Replace with and After that, the time domain inter-class diffusion loss can be obtained respectively ( , ) and frequency domain inter-class diffusion loss ( , ). In order to ensure the consistency of class diffusion in time domain and frequency domain, the present invention also calculates the cross-domain inter-class diffusion loss. In formula (8), the input variable and Replace with and , we can get the cross-domain inter-class diffusion loss ( , ).

[0063] Step 3.3, construct the triple class diffusion loss function;

[0064] right and The intra-class diffusion loss in the time domain can be obtained by using the intra-class diffusion mechanism in step 3.1 and the inter-class diffusion mechanism in step 3.2. ( , ) and time domain inter-class diffusion loss ( , );right and The intra-class diffusion loss in the frequency domain can be obtained by using the intra-class diffusion mechanism in step 3.1 and the inter-class diffusion mechanism in step 3.2. ( , ) and frequency domain inter-class diffusion loss ( , );right and The intra-class diffusion mechanism of step 3.1 and the inter-class diffusion mechanism of step 3.2 can be used to obtain the cross-domain intra-class diffusion loss respectively. ( , ) and cross-domain inter-class diffusion loss ( , ).

[0065] Based on the above intra-class diffusion loss and inter-class diffusion loss in time domain, frequency domain and cross-domain, the present invention constructs a triple-class diffusion loss function:

[0066] (9)

[0067] in, It is the trade-off parameter between intra-class diffusion loss and inter-class diffusion loss in time domain, frequency domain and cross-domain; 、 and Represents the weighted sum of intra-class diffusion loss and inter-class diffusion loss in time domain, frequency domain and cross-domain respectively; express

[0068] 、 and The trade-off parameters of the three. The range is 0.1~1 (1 in this embodiment), The range is 0.1~1 (0.5 in this embodiment).

[0069] Verification results of this example:

[0070] This paper validates the effectiveness of the proposed method through two sets of experiments. The first set of experiments verifies the clustering performance of the time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion on the public UCR dataset. The second set of experiments verifies the clustering performance of the time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion using the WTC weather dataset that includes various weather-related features.

[0071] UCR datasets: UCR stands for the University of California, Riverside. UCR datasets cover a wide range of application scenarios, including biomedical signals, motion recognition, machine condition monitoring, and many other fields. Five of these datasets were selected for experiments. Among them, Fungi is a time-series feature map of fungal images under a microscope; GPOVY is trajectory data from handheld gun movements; Plane is a time-series data map of aircraft engine vibrations; CBF is a simulated dataset defined by local feature extraction using a base library and its application; and Coffee is a coffee dataset based on Fourier transform infrared spectroscopy.

[0072] The WTC dataset (weather-type-classification) is a weather data dataset used for classification tasks. The data input dimensions are 6, including temperature, humidity, wind speed, precipitation, air pressure, and visibility. It includes various weather-related features and categorizes weather into four types: rainy, sunny, cloudy, and snowy.

[0073] The experiment used Normalized Mutual Information and Rand Index as evaluation criteria for diagnostic performance.

[0074] For algorithm comparison, six commonly used time series clustering algorithms were selected for comparison. These algorithms are: Learning Interpretable Representations for Multivariate Time Series Clustering (T-Feat), Convolutional Generative Adversarial Network for Time Series Classification and Clustering (TCGAN), Self-Supervised Time Series Representation Learning via CrossReconstruction Transformer (CRT), Information-Aware Time Series Meta-Contrastive Learning (Infos), Self-Supervised Contrastive Learning for Universal Time Series Representation Learning (T-URL), and Contrastive learning-based multi-view clustering for incomplete multivariate time series (MIMTS).

[0075] The method proposed in this invention, the time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion, and the comparative method were experimented on five UCR datasets and the WTC dataset. The specific clustering results are shown in Tables 1-4.

[0076] Table 1 NMI (normalized mutual information) metrics on the UCR dataset

[0077]

[0078] As can be seen in Table 1, the proposed method achieves an NMI of at least 0.989 on all datasets, significantly outperforming other comparison methods. Furthermore, the proposed method achieves the best NMI results across all datasets, particularly achieving an NMI of 1.0 on datasets such as GPOVY (motion state), Plane (sensor data), CBF (simulated data), and Coffee (spectral signal), fully matching the true labels.

[0079] Table 2 RI (Rand Index) indicators on the UCR dataset

[0080]

[0081] As shown in Table 2, the RI of our method on all datasets is no less than 0.997, significantly outperforming other comparison methods and demonstrating significant overall improvement. Furthermore, our method achieves optimal RI on all datasets, particularly achieving RI of 1.0 on datasets such as GPOVY (motion state), Plane (sensor data), CBF (simulated data), and Coffee (spectral signal), fully matching the true labels.

[0082] Table 3 NMI (normalized mutual information) index on the WTC dataset

[0083]

[0084] As can be seen from Table 3, the NMI value of our method on the WTC dataset is 0.515, which is superior to all the compared methods. This shows that our method can more accurately capture the intrinsic category structure of meteorological data than other compared methods in the clustering task of the WTC weather time series dataset.

[0085] Table 4 RI (Rand Index) indicators on the WTC dataset

[0086]

[0087] As can be seen in Table 4, the RI value of our method on the WTC weather dataset is 0.818, outperforming all the compared methods. The results show that, in the clustering task of the WTC weather time series dataset, our method achieves superior classification consistency for sample pairs compared to the other compared methods.

[0088] The reasons for the superior clustering results in Table 1-4 are: ① The present invention constructs a data enhancement strategy based on mixed multi-granularity perturbations, generates time domain enhancement data in the time domain dimension to retain the local trend and global structure of the time series data, and generates frequency domain enhancement data in the frequency domain dimension to cover the diverse variation patterns of frequency features. ② The present invention designs a multi-domain pattern extractor based on clustering guidance to extract the clustering probability distribution of original data and enhanced data in the time domain and frequency domain, ensuring that data in different domains are comparable in the same space. ③ The present invention constructs a multi-domain reliable class diffusion strategy, and by designing intra-class and inter-class double diffusion mechanisms, constructs a triple class diffusion loss function in the time domain, frequency domain and cross-domain, to promote reliable clustering of time series data. The intra-class diffusion mechanism enhances clustering stability by strengthening the similarity between samples and their nearest reliable neighbors in the latent space, and the inter-class diffusion mechanism maximizes the differences between the representations of each cluster to achieve reliable separation between clusters.

[0089] The above-described embodiments merely express the implementation methods of the present invention, but should not be understood as limiting the scope of the present invention. It should be pointed out that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, which all fall within the scope of protection of the present invention.

Claims

1. A time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion, characterized by: The following steps are involved: Step 1: Construct a data enhancement strategy based on mixed multi-granularity perturbations to generate enhanced data in the time domain and frequency domain; In the time domain, a multi-granularity time domain data enhancement strategy is constructed by combining jitter, scaling, and permutation operations to generate time domain enhanced data that preserves the local trends and global structure of time series data. In the frequency domain, the original time domain data is first converted into original frequency domain data through Fourier transform, and then a multi-granularity frequency domain data enhancement strategy is constructed by combining frequency addition and removal operations to generate frequency domain enhanced data that covers the diverse variation patterns of frequency features. Step 2: Design a cluster-guided multi-domain pattern extractor to extract the clustering probability distribution of the original and enhanced data in the time and frequency domains, ensuring that data from different domains are comparable in the same space. The time-domain pattern extractor uses a bidirectional long short-term memory network and a clustering block cascade architecture, while the frequency-domain pattern extractor uses a three-layer convolutional block and a clustering block cascade architecture. Step 3: Construct a multi-domain reliable class diffusion strategy. By designing a double diffusion mechanism within and between classes, a triple class diffusion loss function in the time domain, frequency domain, and cross-domain is constructed to promote the reliable clustering of time series data. The intra-class diffusion mechanism enhances the clustering stability by strengthening the similarity between samples and their nearest reliable neighbors in the latent space, and the inter-class diffusion mechanism maximizes the differences between the representations of each cluster to achieve reliable separation between clusters.

2. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 1 is characterized in that: The step 1 is specifically as follows: Step 1.1, construct a multi-granularity temporal data enhancement strategy; Given time domain raw data , represents the number of samples, Indicates the number of clusters (categories); express No. Time domain original samples, Indicates the sequence length of the sample; ,enriching its temporal information by constructing a multi-granularity temporal data enhancement strategy; The multi-granularity temporal data augmentation strategy consists of jittering, scaling, and permutation operations; Step 1.2, construct a multi-granularity frequency domain data enhancement strategy; For time domain raw data Each row of time domain original samples adopts the multi-granularity time domain data enhancement strategy of step 1.1 and merges them to obtain time domain enhanced data ; First, the time domain raw data Fast Fourier Transform Get the original frequency domain data ; Then, for The Original frequency domain samples , by constructing a multi-granularity frequency domain data enhancement strategy to generate frequency domain enhanced samples ; The multi-granularity frequency domain data enhancement strategy consists of a local frequency removal operation and a local frequency addition operation.

3. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 2 is characterized in that: In step 1.1, the multi-granularity time domain data enhancement strategy is specifically as follows: The dithering operation adds random noise to the time series data; the scaling operation adds random noise to the original samples in the time domain. Performs element-wise multiplication of and a given scaling factor; the permutation operation first uses a split operation to The time dimension is divided into The multi-granularity time domain data enhancement strategy is as follows: (1), in, express No. Time domain original samples; express Enhanced time domain enhanced samples; Indicates a dither operation; Represents a zoom operation; Represents a permutation operation; Represents random noise, which has a mean of 0 and a variance of Normal distribution ; represents the scaling factor, Obey normal distribution; Represents a split operation; express The number of segments divided; L represents The sequence length of Indicates random shuffling operation; Represents a new time series; Represents the time step index obtained after the permutation operation; Represents an element-wise multiplication operation.

4. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 2 is characterized in that: In step 1.2, the multi-granularity frequency domain data enhancement strategy is specifically as follows: The local frequency removal operation will The mask of the local frequency domain removal operation is multiplied element by element; the local frequency addition operation first generates a random perturbation, and then the random perturbation is multiplied element by element with the mask of the local frequency addition operation; the process of the multi-granularity frequency domain data enhancement strategy is as follows: (2), in, express The Original frequency domain samples; express Enhanced frequency domain enhanced samples; represents the local frequency removal operation; The mask representing the local frequency domain removal operation, The probability of success is Bernoulli distribution ,in represents the probability of retaining the frequency component; represents a local frequency addition operation; represents random disturbance; The mask representing the local frequency addition operation, The probability of success is Bernoulli distribution ; represents a hyperparameter.

5. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 4 is characterized in that: The hyperparameters The range is 0.1~0.

5.

6. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 2 is characterized in that: The step 2 is specifically as follows: Step 2.1, construct a time domain pattern extractor; Time Domain Pattern Extractor Adopting the architecture of bidirectional long short-term memory network and clustering block cascade, yes The bidirectional long short-term memory network block consists of a forward and a backward long short-term memory network in parallel. The forward and backward dependencies of the sequence data are captured based on a bidirectional propagation mechanism. The time-step hidden states of the two are spliced ​​to form a feature representation that integrates the bidirectional context information. The clustering block consists of two fully connected layers and a soft allocation function. The time domain raw data and temporal enhancement data Input to the time domain pattern extractor , get the original probability distribution in time domain and the enhanced probability distribution in time domain; Step 2.2, construct a frequency domain pattern extractor; The frequency domain pattern extractor Adopting the architecture of three-layer convolutional blocks and clustering blocks cascaded, yes The learnable parameters of the convolutional layer are batch normalized, activated by rectified linear units, and then subjected to one-dimensional maximum pooling operations in each convolutional block. The first convolutional block has an additional regularization layer after the pooling layer. The clustering block consists of two fully connected layers and a soft allocation function. The original frequency domain data and frequency domain enhanced data Input to the frequency domain pattern extractor , and obtain the original probability distribution in the frequency domain and the enhanced probability distribution in the frequency domain.

7. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 2 is characterized in that: In step 2: In step 2.1, the time domain pattern extraction process is shown in formula (3): (3), in, and Represent the original probability distribution in time domain and the enhanced probability distribution in time domain respectively; In step 2.2, the frequency domain pattern extraction process is shown in formula (4): (4), in, and represent the original probability distribution in the frequency domain and the enhanced probability distribution in the frequency domain respectively.

8. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 6 is characterized in that: The step 3 is as follows: For time domain raw data With time domain enhanced data Use the time domain pattern extractor in step 2.1 to obtain the original probability distribution in the time domain Time-domain enhanced probability distribution ; For the original frequency domain data With frequency domain enhanced data Use the frequency domain pattern extractor in step 2.2 to obtain the original probability distribution in the frequency domain Frequency domain enhanced probability distribution ; In order to derive reliable time series clustering, based on 、 、 and ,construct a multi-domain reliable cluster diffusion strategy, which realizes reliable clustering of multi-domain data by designing intra-class diffusion mechanism and inter-class diffusion mechanism; Firstly, a reliable intra-class diffusion mechanism is designed for the time domain, frequency domain and cross-domain. By constructing a sample neighborhood set, generating a semantic indicator matrix and designing an intra-class diffusion loss, the reliable similarity between samples and their latent space neighbors is enhanced and the clustering stability is improved. Secondly, a reliable inter-class diffusion mechanism is designed for the time domain, frequency domain and cross-domain. By introducing an inter-class diffusion loss with the goal of maximizing partition entropy, reliable separation between clusters is achieved. Finally, the intra-class diffusion loss and inter-class diffusion loss of the time domain, frequency domain and cross-domain are integrated to construct a triple-class diffusion loss function to achieve stable and reliable clustering.

9. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 8, characterized in that: The details are as follows: Step 3.1, design a reliable intra-class diffusion mechanism; set up ,in and Represent the time domain and frequency domain respectively, and define and Represents domain The original probability distribution and enhanced probability distribution of ; First construct and The similarity matrix between the two , establish cross-view semantic associations between samples through the similarity matrix; then No. Row vector , according to the similarity matrix exist Select the one with the greatest similarity Instances construct their sample neighborhood set ; At the same time, the semantic indicator matrix is ​​introduced ; Among them, the semantic indicator matrix Middle Rank Column element value As shown in formula (5): (5), in, represents the maximum assignment probability index generating function; express No. row vector, express No. row vector; Combine the neighborhood set and the semantic indicator matrix to construct a cross-view neighborhood distillation matrix , providing reliable neighborhood relations for inter-class diffusion; among them, the cross-view neighborhood distillation matrix Middle Rank Column element value As shown in formula (6): (6), Finally, the intra-class diffusion loss of the design domain v is ( , )for: (7), in, ( , )express and Intra-class diffusion loss; is the row normalization operator, is the number of neighbors after distillation, Represents element-by-element multiplication operation, exp represents the exponential function with the natural constant e as the base, express The first Rank Elements of the column, express No. row vector; The formula (7) Replace with and After that, we get the intra-class diffusion loss in the time domain ( , ) and frequency domain intra-class diffusion loss ( , ); while calculating the cross-domain intra-class diffusion loss; in formula (7), the input variable and Replace with and Get cross-domain intra-class diffusion loss ( , ); Step 3.2, design a reliable inter-class diffusion mechanism; Through the domain The original probability distribution and enhanced probability distribution Imposing orthogonal constraints forces the semantic representation to approach the geometric simplex of the semantic space, thereby enhancing the ability to capture boundaries; Will and Each column of is interpreted as a specific cluster representation, and then reliable separation between clusters is achieved by maximizing the difference between the cluster representations; thus, the inter-class diffusion loss with the goal of maximizing partition entropy is introduced: (8), in, express and inter-class diffusion loss; express The mean probability of belonging to the jth cluster in ; express The mean probability of belonging to the jth cluster in ; represents the identity matrix; represents information entropy loss; express No. The clusters represent express No. The clusters represent express and The similarities between the two; The formula (8) Replace with and After that, we get the time domain inter-class diffusion loss ( , ) and frequency domain inter-class diffusion loss ( , ); while calculating the cross-domain inter-class diffusion loss; in formula (8), the input variable and Replace with and , we get the cross-domain inter-class diffusion loss ( , ); Step 3.3, construct the triple class diffusion loss function; right and The intra-class diffusion loss in the time domain is obtained by using the intra-class diffusion mechanism in step 3.1 and the inter-class diffusion mechanism in step 3.

2. ( , ) and time domain inter-class diffusion loss ( , );right and The intra-class diffusion loss in the frequency domain can be obtained by using the intra-class diffusion mechanism in step 3.1 and the inter-class diffusion mechanism in step 3.

2. ( , ) and frequency domain inter-class diffusion loss ( , );right and The intra-class diffusion mechanism of step 3.1 and the inter-class diffusion mechanism of step 3.2 can be used to obtain the cross-domain intra-class diffusion loss respectively. ( , ) and cross-domain inter-class diffusion loss ( , ); Based on the above intra-class diffusion loss and inter-class diffusion loss in time domain, frequency domain and cross-domain, a triple-class diffusion loss function is constructed: (9), in, It is the trade-off parameter between intra-class diffusion loss and inter-class diffusion loss in time domain, frequency domain and cross-domain; 、 and Represents the weighted sum of intra-class diffusion loss and inter-class diffusion loss in time domain, frequency domain and cross-domain respectively; express 、 and The trade-off parameters among the three.

10. The time series clustering method based on multi-granularity enhancement and multi-domain reliable class diffusion according to claim 9, characterized in that: In the formula (9), the hyperparameter The range is 0.1~1, The range is 0.1~1.