High-dimensional time series data classification method and device based on multi-scale diffusion denoising

Through the combination of multi-scale diffusion model and Transformer architecture, the classification and missing data processing of irregular time series are solved, and efficient classification and robustness of irregular time series are achieved.

CN120277526APending Publication Date: 2025-07-08NANJING UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510343774.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-22
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art lacks the task of applying diffusion models to time series classification, interpolation and other tasks, especially in irregular time series, the robustness and classification ability of missing data.

Method used

Using a multi-scale diffusion model method, multi-grained conditional priors are constructed through multi-scale coding, dual attention feature extraction, conditional guidance model and denoising model, and feature extraction and denoising processing are used to perform feature extraction and denoising processing to realize the classification of irregular time series.

Benefits of technology

It improves the classification ability of irregular time series and the robustness of missing data, enhances the coding ability and data feature capture ability of the model, and provides accurate estimation of the real label probability distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277526A_ABST
    Figure CN120277526A_ABST
Patent Text Reader

Abstract

The invention discloses a high-dimensional time series data classification method and device based on multi-scale diffusion denoising. According to the method, a multi-scale condition guidance strategy is introduced to guide the denoising process of time sequence labels, and condition priori of different scales is used in the diffusion process to adjust each step. The method comprises the following steps of: encoding time sequence data from two angles of variable and time, extracting features, obtaining multi-scale feature representation of a time sequence, mining a dependency relationship of the sequence on a time dimension and a variable dimension by using a double attention mechanism, and generating condition priori on different scales; a standard denoising diffusion probability model is used for training a process, a fusion vector in a potential space is obtained by using a multi-scale encoder, then the fusion vector and a time step are embedded, and a full connection layer is used for predicting noise. According to the method, the result accuracy of the high-dimensional multivariable time series in the classification task and the robustness of processing irregular time series data containing missing values are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data processing, and particularly to a method and device for classifying time series data in a time series data mining task. Background Art

[0002] Irregular time series widely exist in fields such as medical monitoring, financial transactions, Internet of Things data, and meteorological observations. Compared with traditional regular time series, the data points are unevenly distributed on the time axis, and there may be missing, asynchronous sampling, and dynamically changing time intervals. This characteristic poses great challenges to data modeling, prediction, and analysis. In recent years, technologies such as deep learning, Bayesian modeling, and graph neural networks have provided new ideas for modeling irregular time series. In particular, methods based on Neural Ordinary Differential Equations (Neural ODEs), Temporal Transformers, and Diffusion Models can effectively capture the impact of time interval changes on sequence patterns.

[0003] Irregularly sampled multivariate time series exhibit two prominent characteristics: within-sequence irregularity and between-sequence differences. Within-sequence irregularity means that in practice, the variable signals within each time series are usually recorded at irregular intervals, breaking the assumption of traditional regularly sampled time series. In addition, between-sequence differences indicate that there are significantly different sampling rates between multiple time series, resulting in a significant imbalance in the appearance of different signals. These two data characteristics pose great challenges to modeling irregularly sampled multivariate time series. How to improve prediction accuracy, enhance data robustness, and solve missing data filling and cross-time scale fusion has become a key issue.

[0004] Recently, the Denoising Diffusion Probabilistic Model (DDPM) has achieved excellent results in image generation and synthesis tasks by iteratively improving the quality of a given image. Specifically, DDPM is a generative model based on a Markov chain that models the data distribution by simulating a diffusion process that evolves the input data towards the target distribution. Although some pioneer works have attempted to use diffusion models for image segmentation and object detection tasks, their potential in time series has not been fully explored. There are some works that use diffusion models to impute irregular time series with missing values, but few works directly study the performance of diffusion models in time series classification tasks. Summary of the Invention

[0005] In view of the deficiency in the prior art that there is a lack of applying diffusion models to various downstream tasks of time series, such as classification, interpolation and other tasks, the present invention provides a multivariate time series classification method and device based on a multi-scale diffusion model, which improves the classification ability for irregular time series and the robustness in the presence of missing data.

[0006] In order to achieve the above invention purpose, the technical solution of the present invention is as follows:

[0007] A high-dimensional time series data classification method based on multi-scale diffusion denoising, comprising the following steps:

[0008] For an irregular time series data instance, encode it based on time points and variables, construct time series of different scales, and use upsampling and downsampling to unify the time series of each scale to the same length;

[0009] For time series of different scales, use a dual attention feature extraction module to extract features in the time and variable dimensions respectively by using the attention mechanism;

[0010] Fuse the features of dual attention, use a conditional guidance model to predict the corresponding label, construct a conditional prior at each scale, and calculate the cross-entropy loss between the label predicted by the conditional guidance model and the true label value to train the conditional guidance model;

[0011] Add noise to the true labels of the training set, fuse the time step embedding and the conditional prior constructed by using the conditional guidance model, and predict the denoised true labels based on the fused vector, and weight the losses on different conditional priors to train the denoising model;

[0012] Use the trained conditional guidance model to obtain the multi-granularity conditional priors of the test set, and then use the trained denoising model to remove the noise in the labels to generate the final prediction.

[0013] Further, encoding based on time points and variables includes:

[0014] For a given irregular time series data set where x i represents the multivariate data of the i-th irregular time series containing missing values, y i represents the true label of the i-th time series, P is the total number of time series, and the multivariate data of each time series is expressed as where K represents the number of variables, L k represents the number of observed values of the k-th variable, that is, the number of time points, represents the timestamp and variable value of the j-th observed value; construct a time step matrix T ∈ R L and a value matrix X ∈ R K×L, the variable matrix E ∈ R K×L , and then perform embedding to construct the input data matrix H ∈ R K×L×D , where D represents the dimension of the embedding vector.

[0015] Furthermore, the attention mechanism is adopted to extract features in the time and variable dimensions for time series of different scales, including:

[0016] At each scale, the dynamic time warping algorithm is used to select the L time-point sequences that are closest to the L time-point sequences of the original time series to form a new sequence, and the transformation matrix A i is used to transform the input data matrix H i into the latent space feature representation Z i : i

[0017] The attention mechanism of the Transformer architecture is used to capture the attention features of Z i in the time and variable dimensions, including: for the feature vectors in the hidden state, first split the vector into K data sequences of L×D, apply the self-attention mechanism to extract the sequence features in the variable dimension, and then split the vector into L data sequences of K×D, apply the self-attention mechanism to extract the sequence features in the time dimension.

[0018] Furthermore, the conditional guidance model is a classifier using a fully connected layer.

[0019] Furthermore, the denoising model follows the design of the standard denoising diffusion probabilistic model and is divided into two stages: the forward diffusion stage, i.e., the training stage, and the reverse diffusion stage, i.e., the inference stage. In the forward diffusion stage, the true response variable y0 adds Gaussian noise through a diffusion process conditional on the time step t, and this diffusion process samples from the uniform distribution of [1, T]. The TransUNet is used to implement the denoising network to parameterize the reverse diffusion process and learn the noise distribution in the forward process; in the reverse diffusion stage, the trained TransUNet (∈ θ ) generates the final prediction by converting the noise variable distribution p θ (y T ) into the true distribution p θ (y0) where θ represents the parameters of the denoising network, ∈ θ is the output predicted by the denoising model, y T represents the label value after adding noise for T steps, and y0 represents the original label value.

[0020] Furthermore, the denoising model executes the conditional generation method of denoising diffusion, including:

[0021] Given an input irregular time series dataset Obtain the time series feature embedding ρ(x) through multi-scale encoding and the multi-scale conditional guidance model to generate priors at 3 different scales Then for y0, Perform t steps of noise addition to generate the noise variable y t , Combine the 4 noise variables y t , with their respective priors and project them into the latent space separately; then integrate the 4 projected embeddings with the feature embedding ρ(x) in the denoising network TransUNet respectively to predict the noise distribution sampled for y t , :

[0022]

[0023] where θ are the parameters of the denoising network TransUNet, N(·,·) represents the Gaussian distribution, and I is the identity matrix.

[0024] Furthermore, in the training stage of the denoising model, apply the diffusion process to the true label y0 and different priors to generate 4 noise variables, where:

[0025]

[0026] where, ∈~Ν(0,Ι), α t = 1-β t , {β t} t=1:T ∈(0,1) T is the noise coefficient at each step. After that, input the concatenated vector of the noise variable y t and the multi-scale prior into the denoising network TransUNet to estimate the noise distribution, and its formula is:

[0027]

[0028] where f(·) represents the projection layer to the latent space, [·] is the concatenation operation, E(·) and D(·) are the encoder and decoder of TransUNet. In the forward process, minimize the noise estimation loss:

[0029]

[0030] When calculating the loss, use the mean squared error noise estimation loss for the predicted noise of y t to train, for The maximum mean discrepancy is used to measure the loss, and weights are assigned to jointly train the entire denoising network.

[0031] A high-dimensional time series data classification device based on multi-scale diffusion denoising, comprising:

[0032] A preprocessing module, which is used for an irregular time series data instance to be encoded based on time points and variables, construct time series of different scales, and use upsampling and downsampling to unify the time series of each scale to the same length;

[0033] A multi-scale feature extraction module, which is used for time series of different scales, and adopts a dual attention feature extraction module to extract features in the time and variable dimensions respectively by using the attention mechanism;

[0034] A conditional prior construction module, which is used to fuse the features of dual attention, use the conditional guidance model to predict the corresponding labels, construct the conditional prior at each scale, and calculate the cross-entropy loss between the labels predicted by the conditional guidance model and the true label values to train the conditional guidance model;

[0035] A denoising model training module, which is used to add noise to the true labels of the training set, fuse the time step embedding and the conditional prior constructed by using the conditional guidance model, and predict the denoised true labels based on the fused vector, and weight the losses on different conditional priors to train the denoising model;

[0036] A true label prediction module, which is used to obtain the multi-granularity conditional priors of the test set by using the trained conditional guidance model, and then use the trained denoising model to remove the noise in the labels to generate the final prediction.

[0037] The present invention also provides an electronic device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the program is executed by the processor, the steps of the high-dimensional time series data classification method based on multi-scale diffusion denoising as described above are implemented.

[0038] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the high-dimensional time series data classification method based on multi-scale diffusion denoising as described above are implemented.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] (1) Emphasize the importance of intra - sequence irregularity and inter - sequence differences in irregularly sampled multivariate time series, calculate the attention mechanism from two perspectives of the time dimension and the variable dimension for feature extraction, which is beneficial to providing richer data information for the model. (2) Extract data information from a multi - scale perspective to form a multi - granular conditional guidance scheme, which is beneficial to improving the robustness of the model, better capturing data features, and enhancing the encoding ability. (3) Propose a new model based on denoising diffusion, introduce the relatively advanced diffusion model into the processing of time - series data, which is beneficial to better exerting the powerful strength of the diffusion model, and can accurately classify irregular time - series data, laying a foundation for the wide application of diffusion models in time series in the future. (4) Inject covariate dependence and a pre - trained conditional guidance classifier into the forward and reverse diffusion chains to construct a denoising diffusion probability model, providing an accurate estimate of the true label probability distribution. Description of the Drawings

[0041] Figure 1 is the flowchart of the multi - variable time - series classification method based on the multi - scale diffusion model;

[0042] Figure 2 is the schematic diagram of the model for multi - variable time - series classification based on the multi - scale diffusion model. Detailed Embodiment

[0043] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments and the corresponding drawings. At the same time, it should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent modifications of the present invention all fall within the scope defined by the appended claims of this application.

[0044] Refer to Figure 1 and Figure 2 In the first embodiment of the present invention, a multi - variable time - series classification method based on a multi - scale diffusion model is proposed. The method is completed through a pre - trained conditional guidance classifier and a conditional generation model based on denoising diffusion. Among them, the conditional guidance classifier includes a multi - scale encoding module and a dual - attention feature extraction module. In the context of the present invention, multi - variable time series and high - dimensional time series have the same meaning. Specifically, the method includes the following steps:

[0045] (1) Multi - scale encoding step: For an irregular time - series data instance, encode based on time points and variables, construct time series of different scales, and use upsampling and downsampling to unify different scales for convenient subsequent model processing.

[0046] (2) Dual attention feature extraction step: For time series of different scales, the attention mechanism is adopted to extract features in the time and variable dimensions.

[0047] (3) Conditional guidance classification step: Integrate the features of dual attention to construct the conditional prior at each scale. Calculate the cross-entropy loss based on the labels predicted by the conditional guidance model and the true label values, and train the conditional guidance model.

[0048] (4) Conditional denoising step: Add noise to the true labels of the training set, integrate the time step embedding and the conditional prior, predict the denoised true labels, weight the losses on different conditional priors, and train the denoising model.

[0049] (5) True label prediction step: Use the trained conditional guidance model to obtain the multi-granularity priors of the test set, and then use the trained denoising model to remove the noise in the labels to generate the final prediction.

[0050] In the embodiments of the present invention, taking the ICU data of patients in the medical monitoring field as an example, the specific implementation of each step is described in detail. It should be understood that this method is also applicable to time series in other fields.

[0051] According to the embodiments of the present invention, in the multi-scale encoding step, for a given irregular time series data set where x i represents the ICU data of the i-th patient in 48 hours, y i represents whether the i-th patient finally survives, taking values of 0 or 1, and P is the total number of patients. The ICU data of each patient can be expressed as where K represents the number of variables, L k represents the number of observed values of the k-th variable, that is, the number of time points, which is also the length of the time series, represents the timestamp and variable value of the j-th observed value. Construct the time step matrix T∈R L , the value matrix X∈R K×L , the variable matrix E∈R K×L , and then further perform embedding to construct the input data matrix H∈R K×L×D :

[0052] H = f val (X) + f type (E) + f time (T)

[0053] where D represents the dimension of the embedding vector. f val represents a fully connected mapping used to convert a single value into an embedding of size D, and f type represents a lookup table that converts categorical variable indicators into type embeddings (also of size D), and ftime is a function for obtaining time-related embeddings. In irregular time series, since some metrics are recorded irregularly, the time step lengths of different user sequences are different, such as L = 128 or L = 156. When dealing with irregular time series, it is often necessary to interpolate data, fill in missing values, or handle time inconsistency issues through models. In the present invention, by constructing transformation tensors of different shapes, the time series are unified to the same scale. For example, for each time series H of different lengths, 32 most representative time points are sampled using the Dynamic Time Warping (DTW) algorithm, and the tensor shape is unified into [batch_size, variates, timesteps, embedding_size] to facilitate the processing of subsequent modules in the model. batch_size is the number of samples used in one training process. In deep learning, training data is usually divided into multiple small batches (batches) for processing, and the batch size represents the number of samples in each batch. variates refers to the number of features or different variables contained in each sample. In time series or multi-channel data, variates represent different attributes or observation dimensions in the data sample. timesteps is the length of the time series data, representing the number of time points in the data. Each time step corresponds to an observation value or feature value in the data. In sequence modeling tasks, timesteps represents the number of time or events in the input sequence. embedding_size is the dimension of the embedding space, usually used to represent the dimension of the feature representation obtained through a certain encoding (such as word embedding or position embedding).

[0054] Specifically, after embedding the time series data to obtain H, a multi-scale encoder is used to extract multi-scale features from the time series data. For the original time series length L, 3 time series with scales scale1, scale2, and scale3 are constructed by selecting different ratios such as L / 4, L / 2, and L. At each scale scale i using the DTW algorithm, the L time point sequences that are closest to the L time point sequences of the original time series are selected to form a new sequence. To construct the new sequence, a transformation matrix is needed i to convert the original data representation into a latent space feature representation where n refers to the nth layer of the multi-scale encoder. For example, if 3 scales are set, then the value of n is 1, 2, 3.

[0055] ​In the dual attention feature extraction step, the attention mechanism of the Transformer architecture is used to learn high-level representations from the unified input data (denoted as Z obtained from multi-scale feature extraction) separately in the variable dimension and the time dimension. The attention formula is as follows:

[0056]

[0057] where Q is the query statement, K is the keyword, V is the value, and d k is equal to the length of the word vector. First, this method applies the attention mechanism from the time dimension, splits the vector Z into K data sequences of L×D, and captures the temporal relationships within the sequence at this scale z t :

[0058] z t = Attention(Z i ,Z i ,Z i )

[0059] Attention(Z i ,Z i ,Z i ) represents calculating the self-attention of Z i .

[0060] After adding the residual connection to the original input, the vector is split into L data sequences of K×D, and the tensor shape is changed from [batch size, variates, timesteps, embedding size] to [batch size, timesteps, variates, embedding size]. Then, the attention mechanism is applied from the variable dimension to capture the relationships between variables in different sequences z v .

[0061] Z' i = Rearrange(Z i + z t )

[0062] z v = Attention(Z' i ,Z' i ,Z' i )

[0063] In the conditional guidance classification step, the dual-dimensional features are added together and mapped to the latent space to obtain a new H iInput the next scale for calculation. After the entire model obtains the features at each scale, the features are aggregated and then input into the decoder. The decoder consists of a fully connected network. After passing through the fully connected network, the probability distributions of different classes are output, and then the cross-entropy loss is calculated between the predicted probability distributions and the true 0 / 1 labels. The conditional prior model is pre-trained to make the accuracy of the predicted labels reach about 90%.

[0064] In the conditional denoising step, different from the unconditional diffusion model, the conditional diffusion model learns to input random noise and conditional information under a given condition, and learns how to denoise and recover the data. The generated samples depend on the additional conditional information. In this work, the diffusion process is as follows:

[0065]

[0066] where is the prior knowledge of the time series data x and the true label y0, and y T is the noisy data after T steps of diffusion, and I is the identity matrix. Use the pre-trained network, i.e., the pre-trained conditional guidance classifier mentioned above, to approximate E[y|x]. For all time steps t, the diffusion step from adding noise at t - 1 to t is:

[0067]

[0068] For any time step t, the diffusion step from 0 to t is:

[0069]

[0070] where α t := 1 - β t , the growth scale of the noise {β t} t=1:T ∈(0, 1) T . Correspondingly, the posterior of the forward process:

[0071]

[0072] where is the denoising framework to be constructed below with y t , y0, as the input.

[0073] In the scenario of this paper, given the input irregular time series data set Input it into the pre-trained conditional guidance classifier to generate priors at different scales Add noise to y0, respectively to generate 4 noise variables y t , Then, four noise variables y t , and their respective priors are combined and projected into the latent space separately. Further, the three projection embeddings are integrated with the temporal feature embedding ρ(x) in the denoising TransUNet respectively, and the noise distribution sampled for y t , is predicted:

[0074]

[0075] where θ are the parameters of the denoising TransUNet, and the denoising TransUNet refers to the denoising module in Figure 2 , and N(·,·) represents the Gaussian distribution, and I is the identity matrix.

[0076] Traditional denoising models use the U-Net architecture to complete image segmentation tasks. However, due to the inherent locality of convolutional operations, they usually show limitations in modeling explicit long-range relationships. Therefore, these architectures usually have poor performance. When facing time series tasks, the Transformer designed for sequence-to-sequence prediction has become an alternative architecture with an innate global self-attention mechanism. However, due to the lack of low-level details, it may lead to limited localization ability. The Transformer treats the input as a 1D sequence and focuses on modeling the global context at all stages, resulting in low-resolution features lacking detailed localization information. The TransUNet combines the advantages of the Transformer and the U-Net. On the one hand, the Transformer encodes the tokenized image patches from the convolutional neural network (CNN) feature map into an input sequence for extracting the global context. On the other hand, the decoder upsamples the encoded features and then combines them with the high-resolution CNN feature map for precise localization. The combination of the Transformer and the U-Net can enhance finer details by restoring local spatial information. To make up for the loss of feature resolution brought by the Transformers, the TransUNet adopts a hybrid CNN-Transformer architecture, which can utilize both the CNN to obtain detailed high-resolution spatial information of the features and the Transformers to obtain the encoded global context information.

[0077] The present invention is a conditional generation model based on denoising diffusion, following the design of DDPM, which is mainly divided into two stages: the forward diffusion stage (training) and the reverse diffusion stage (inference). In the forward process, the true response variable y0 is added Gaussian noise through a diffusion process conditioned on the time step t, which is sampled from a uniform distribution of [1, T]. TransUNet is used as the denoising network to parameterize the reverse diffusion process, and the noise distribution is learned in the forward process. In the reverse diffusion process, the trained TransUNet (∈ θ ) generates the final prediction by transforming the noise variable distribution into the true distribution.

[0078] Whether it is U-Net or TransUNet, the main object to be processed is images. In order to obtain the feature embedding ρ(x), the existing method is to use ResNet as the encoder, which includes 1 input layer and 5 convolutional layers. Each convolutional layer has multiple convolutional kernels of different sizes to deeply extract image features. If the object is converted to a time series, the convolutional method is not sufficient to extract the position features and global features of the time series. Therefore, when using TransUNet, its structure needs to be adjusted to meet the requirements of the time series conditional diffusion model. Since a multi-scale encoder is used in the pre-training, in the encoder module of the denoising model, the pre-trained multi-scale encoder is still used. In order to connect to the subsequent TransUNet module, a high-dimensional linear layer needs to be passed through the multi-scale encoder module to obtain the fusion vector in the latent space, and the shape of this vector is determined by the shape of the feature vector output by the encoder.

[0079] In ResNet, the input original data is an image dataset with the shape of [batch_size, channels, height, width]. After passing through 1 input layer and 5 convolutional layers, a high-dimensional vector is generated, such as setting the shape to [batch_size, 4096], and a linear layer with an output dimension of 4096. In the time series multi-scale encoder, the input original data is a time series dataset with the shape of [batch_size, variates, time_steps, embedding_size]. Different from images, for time series, the output dimension can generally be adjusted to a smaller value, such as 128 or 256. In order to adjust the response embedding at the time step of the diffusion process, a Hadamard product needs to be performed between the fusion vector and the time step embedding. Here, the main inputs that the diffusion model needs to process are y t , ρ(x) and the time step t. The output is the prediction of the denoising model from y t-1 step to y tThe Gaussian noise added in the first step. In the present invention, a multi-scale encoder is used to output the temporal feature embedding ρ(x), and Gaussian noise is added to the original label for t steps to obtain y t , for a specified time step (usually 500, 1000, etc.) for embedding, and the shape of the embedding vector is the same as that of the temporal feature embedding ρ(x), which is convenient for subsequent vector fusion operations.

[0080] To adjust the response embedding over the time step, this method performs a Hadamard product between y t and the time step t embedding. Then, another Hadamard product is performed on the temporal feature embedding ρ(x) to integrate the three inputs. The Hadamard product is an element-wise multiplication operation. In image processing, the Hadamard product can be used to fuse the features of different images or weight certain regions so that the features of important regions receive more attention. Some networks can control the interaction and weighting between features element-wise through the Hadamard product. This way enables the model to adaptively adjust the feature representation, thereby improving performance. The vector output after the Hadamard product passes through two consecutive fully connected layers, and each layer is followed by a Hadamard product with the time step t embedding. Except for the output layer, all fully connected layers are accompanied by a batch normalization layer and a Softplus non-linear layer. For the conditional guidance classifier This method uses the standard cross-entropy loss as the objective of the model. After pre-training the model for 40 epochs, the denoising diffusion model and the model are jointly trained to complete the end-to-end prediction task.

[0081] In the training stage of the denoising model, the diffusion process is applied to the true label y0 and different priors to generate 4 noise variables, where:

[0082]

[0083] where, ∈~Ν(0,Ι), α t =1-β t , {β t} t=1:T ∈(0,1) T is the noise coefficient for each step. Then, the concatenated vector of the noise variable y t and the prior is input into the denoising model TransUNet(∈ θ ) to estimate the noise distribution: In this method, ∈ is a noise value with the same shape as the true label y0, conforming to the Gaussian distribution, and used to be superimposed on the true label y0. ∈ θis the output of the model, i.e., the ∈ value predicted by the denoising module in this method, while θ is the parameter of the deep neural network in the denoising module.

[0084]

[0085] Among them, f(·) represents the projection layer to the latent space. [·] is the concatenation operation. E(·) and D(·) are the encoder and decoder of TransUNet. The time series multi-scale feature embedding ρ(x) is further integrated with the projected noise embedding in TransUNet, enabling the model to obtain a more robust feature representation. In the forward process, the noise estimation loss is minimized:

[0086]

[0087] When calculating the loss, the mean squared error (MSE) noise estimation loss is used for the predicted noise of y t to co-train, and for the maximum mean discrepancy (MMD) is used to measure the loss, and then they are weighted and added together.

[0088] In the inference stage, i.e., the true label prediction step, for the given irregular time series dataset X, where X represents the patient's indicator data and Y represents the patient's survival label, first it is input into the conditional guidance classifier to obtain the prior Then, following the DDPM process, the trained TransUNet determined by and the time series feature embedding ρ(x) is used to iteratively denoise from the random prediction y T to finally predict

[0089] In actual implementation, common deep learning frameworks such as Tensorflow, Pytorch, etc. can be used to implement the method of the present invention.

[0090] The second embodiment of the present invention provides a high-dimensional time series data classification device based on multi-scale diffusion denoising, including:

[0091] A preprocessing module, which is used for an irregular time series data instance to be encoded based on time points and variables, construct time series of different scales, and unify the time series of each scale to the same length by using upsampling and downsampling;

[0092] A multi-scale feature extraction module, which is used for time series of different scales, and a dual attention feature extraction module is adopted to respectively extract features in the time and variable dimensions by using the attention mechanism;

[0093] The conditional prior construction module is used to fuse the features of dual attention, use the condition to guide the model to predict the corresponding label, construct the conditional prior at each scale, and calculate the cross-entropy loss between the label predicted by the condition-guided model and the true label value to train the condition-guided model;

[0094] The denoising model training module is used to add noise to the true label of the training set, fuse the time-step embedding and the conditional prior constructed by the condition-guided model, and predict the denoised true label based on the fused vector, and weight the losses on different conditional priors to train the denoising model;

[0095] The true label prediction module is used to obtain the multi-granularity conditional prior of the test set by using the trained condition-guided model, and then use the trained denoising model to remove the noise in the label to generate the final prediction.

[0096] It should be understood that the high-dimensional time series data classification device based on multi-scale diffusion denoising in the embodiments of the present invention can implement all the technical solutions in the above method embodiments. The functions of its respective functional modules can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions in the above embodiments, which will not be elaborated here.

[0097] The present invention also provides an electronic device, including: one or more processors; a memory; and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the program is executed by the processor, it implements the steps of the above-mentioned high-dimensional time series data classification method based on multi-scale diffusion denoising.

[0098] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above-mentioned high-dimensional time series data classification method based on multi-scale diffusion denoising.

[0099] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, a computer device, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0100] The present invention will be described with reference to the flowchart of a method according to an embodiment of the present invention. It should be understood that each process in the flowchart and the combination of processes in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 or multiple processes. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 or multiple processes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 or multiple processes.

Claims

1. A classification method for high-dimensional time series data based on multi-scale diffusion denoising, characterized in that, Including the following steps: For an irregular time series data instance, encode based on time points and variables, construct time series at different scales, and use upsampling and downsampling to unify the time series at each scale to the same length; For time series at different scales, use a dual attention feature extraction module to extract features in the time and variable dimensions respectively using the attention mechanism; Fuse the features of dual attention, use the conditional guidance model to predict the corresponding labels, construct the conditional prior at each scale, and calculate the cross-entropy loss between the labels predicted by the conditional guidance model and the true label values to train the conditional guidance model; Add noise to the true labels of the training set, fuse the time step embedding and the conditional prior constructed using the conditional guidance model, and predict the denoised true labels based on the fused vector, weight the losses on different conditional priors, and train the denoising model; Use the trained conditional guidance model to obtain the multi-granularity conditional priors of the test set, and then use the trained denoising model to remove the noise in the labels to generate the final prediction.

2. The method according to claim 1, wherein Encoding based on time points and variables includes: For a given irregular time series dataset where x i represents the multivariate data of the i-th irregular time series containing missing values, and y i represents the true label of the i-th time series. P is the total number of time series. The multivariate data of each time series is represented as where K represents the number of variables, and L k represents the number of observations of the k-th variable, i.e., the number of time points, represents the timestamp and variable value of the j-th observation; construct the time step matrix T ∈ R K , the value matrix X ∈ R K×K , the variable matrix E ∈ R K×K , and then perform embedding to construct the input data matrix H ∈ R K×K×D , where D represents the dimension of the embedding vector.

3. The method according to claim 2, wherein Using the attention mechanism to extract features in the time and variable dimensions for time series at different scales, including: At each scale, the dynamic time warping algorithm is used to select the L time-point sequences that are closest to the L time-point sequences of the original time series to form a new sequence. Using the transformation matrix A i transform the input data matrix H i into the latent space feature representation Z i : i Capture Z using the attention mechanism of the Transformer architecture i Attention features in the time and variable dimensions, including: for the feature vectors in the hidden state, first split the vectors into K data sequences of L×D, apply the self-attention mechanism to extract the sequence features in the variable dimension, and then split the vectors into L data sequences of K×D, apply the self-attention mechanism to extract the sequence features in the time dimension.

4. The method according to claim 1, wherein The conditional guidance model is a classifier using fully connected layers.

5. The method according to claim 1, characterized in that The denoising model follows the design of the standard denoising diffusion probabilistic model and is divided into two stages: the forward diffusion stage, which is the training stage, and the reverse diffusion stage, which is the inference stage. In the forward diffusion stage, the true response variable y0 is added Gaussian noise through a diffusion process conditioned on the time step t, which is sampled from a uniform distribution of [1, T]. The TransUNet is used to implement the denoising network to parameterize the reverse diffusion process, and the noise distribution is learned during the forward process; in the reverse diffusion stage, the trained TransUNet (∈ θ ) generates the final prediction by transforming the noise variable distribution p θ (y T ) into the true distribution p θ (y0) where θ represents the parameters of the denoising network, ∈ θ is the output predicted by the denoising model, y T represents the label value after adding noise in T steps, and y0 represents the original label value.

6. The method according to claim 5, characterized in that, The denoising model performs a conditional generation method for denoising diffusion, including: Given an input irregular time series dataset Obtain the time series feature embedding ρ(x) through multi-scale encoding and a multi-scale conditional guidance model to generate priors at three different scales Then for Perform t steps of noise addition to generate noise variables Combine the four noise variables with their respective priors and project them into the latent space separately; then integrate the four projected embeddings with the feature embedding ρ(x) in the denoising network TransUNet to predict the noise distribution for the sampled noise: Where θ is the parameter of the denoising network TransUNet, N(·,·) represents the Gaussian distribution, and I is the identity matrix.

7. The method according to claim 6, characterized in that In the training stage of the denoising model, apply the diffusion process to the true label y0 and different priors to generate 4 noise variables, where: where, ∈~Ν(0,Ι), α t = 1-β t , {β t} t=1:T ∈(0,1) T is the noise coefficient at each step. After that, the noise variable y t and the connection vector of the multi-scale prior are input into the denoising network TransUNet to estimate the noise distribution, and its formula is: Where f(·) represents the projection layer to the latent space, [·] is the concatenation operation, E(·) and D(·) are the encoder and decoder of TransUNet, and in the forward process, minimize the noise estimation loss: When calculating the loss, the predicted noise of y t is trained using the noise estimation loss of the mean squared error. For the maximum mean discrepancy is used to measure the loss and weights are assigned to co-train the entire denoising network.

8. A high-dimensional time series data classification device based on multi-scale diffusion denoising, characterized in that, Including: A preprocessing module for encoding an irregular time series data instance based on time points and variables, constructing time series at different scales, and using upsampling and downsampling to unify the time series at each scale to the same length; A multi-scale feature extraction module for using a dual attention feature extraction module to extract features in the time and variable dimensions respectively for time series at different scales using the attention mechanism; A conditional prior construction module for fusing the features of dual attention, using the conditional guidance model to predict the corresponding labels, constructing the conditional prior at each scale, and calculating the cross-entropy loss between the labels predicted by the conditional guidance model and the true label values to train the conditional guidance model; A denoising model training module for adding noise to the true labels of the training set, fusing the time step embedding and the conditional prior constructed using the conditional guidance model, and predicting the denoised true labels based on the fused vector, weighting the losses on different conditional priors, and training the denoising model; A true label prediction module for using the trained conditional guidance model to obtain the multi-granularity conditional priors of the test set, and then using the trained denoising model to remove the noise in the labels to generate the final prediction.

9. An electronic device, characterized in that, Including: One or more processors; A memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the programs are executed by the processor, the steps of the high-dimensional time series data classification method based on multi-scale diffusion denoising according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the high-dimensional time series data classification method based on multi-scale diffusion denoising according to any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Urban activity prediction method and system based on Transform architecture and denoising diffusion probability model

    CN120471108A

  • Urban activity prediction method and system based on Transformer architecture and denoising diffusion probability model

    CN120471108B

  • Water work order quantity prediction method based on diffusion model

    CN120849879A