Classification method, system and device for irregular time series
By building a classification model containing time-aware interpolation and offset selection modules, the existing interpolation methods are solved, and efficient and accurate data interpolation and classification performance are improved.
Patent Information
- Application Number
- CN202510246949.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing interpolation methods based on irregular time series have problems such as insufficient accuracy, high calculation cost, and inflexible data processing methods. Especially when dealing with multivariate time series, traditional methods ignore the nonlinear and dynamic changes of the data, resulting in inaccurate filling results and large calculation overhead. The unified time point selection strategy cannot adapt to the change patterns of different characteristics, which may lead to the loss of important information.
A classification model including a time-aware interpolation, a front dual self-attention module, an offset selection module, a rear dual self-attention module and a classifier is constructed. This model captures the dynamic relationship of the time series through a time-aware attention mechanism, and adaptively selects the most critical data points with the offset selection strategy to ensure that the classification model can make full use of data characteristics and improve the interpolation quality and classification performance.
It significantly improves the quality and computing efficiency of data interpolation, avoids information loss, simplifies time series representation, reduces computational complexity and resource consumption, and improves classification performance and model generalization capabilities.
Smart Images

Figure CN119740148B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data processing, and in particular relates to a classification method suitable for irregular time series and a corresponding system and device. Background Art
[0002] In the fields of science and engineering, especially in healthcare, biomechanics, climate science, and environmental monitoring, the collection of time series data often faces the problems of uneven sampling and missing data. The root causes of these problems are diverse, including limitations in sensor accuracy, equipment failure, unstable network transmission, and changes in the external environment. For example, in the field of healthcare, real-time monitoring of multiple physiological indicators of patients (such as heart rate, blood pressure, blood oxygen saturation, respiratory rate, etc.) usually relies on high-precision sensors and equipment. However, due to equipment failure, sensor failure, data transmission delay or interference, data may be missing or recorded incorrectly, thus affecting subsequent analysis and decision-making. In environments such as ICUs (intensive care units), where patients' vital signs change rapidly and diversely, the problem of uneven data sampling is particularly prominent. Especially when the patient's condition changes suddenly, medical equipment may not be able to provide complete multi-dimensional data at each sampling moment, resulting in missing data, which brings great challenges to data processing and analysis. Therefore, handling missing values is a major problem in the classification of irregular time series.
[0003] To address this problem, existing literature mainly uses interpolation methods to fill missing data. Traditional interpolation techniques, such as forward filling, zero interpolation, and average interpolation, are widely used to process irregular time series due to their simplicity, ease of implementation, and low computational cost. These methods usually fill missing data by directly using the values of adjacent known data points or other simple rules (zero value, mean, median, etc.) to maintain the continuity and integrity of the data, thereby reducing the impact of missing data on subsequent analysis. However, these simple interpolation methods may introduce large errors when processing more complex time series data, especially when data is severely missing or the data changes dramatically. In order to improve the quality of interpolation, researchers have also explored interpolation methods based on machine learning and deep learning, which can better capture the temporal characteristics of data and provide more accurate interpolation results. For example, the Score-CDM designed by Shunyang Zhang et al. is a fractionally weighted convolutional diffusion model specifically used for interpolation of multivariate time series data. By combining the fractionally weighted convolution module and the adaptive receiving module, the model can effectively balance local and global temporal features, thereby improving the data completion effect. MissNet developed by Kohei Obata et al. is based on a state-space model, using time dependency and mining the interrelationships between sequences by switching sparse networks, and successfully implements missing value interpolation. The ReCTSi method proposed by Zhichen Lai et al. provides an efficient interpolation solution in resource-constrained environments by decoupling pattern learning and integrity-aware attention mechanisms. The TSDE method developed by Zineb Senane et al., which combines self-supervised learning, diffusion process, and IIF masking strategy, can effectively handle missing values in multivariate time series data and improve interpolation performance by learning a common time series representation. Compared with traditional interpolation methods, neural network-based techniques can better capture complex patterns and potential nonlinear relationships in data, thereby providing more accurate and robust interpolation results. Although these methods perform well in improving interpolation quality, they are also accompanied by high computational costs. In addition, simply combining the interpolation model with the classifier directly may lead to a lack of effectiveness in the interpolation results, which cannot fully support the classification task, thereby affecting the ideal performance of the classification effect.
[0004] Recently, a new idea has been proposed to further screen and process the imputed values after the imputation is completed to improve the final analysis and classification performance. For example, Ane Blázquez-García proposed a selective interpolation method that can identify the missing time step subset that is most needed to be estimated in multivariate time series data. This method not only effectively reduces the estimation error, but also maximizes the retention of the characteristics of the original data, thereby simplifying the time series and reducing the computational complexity. Subsequently, Chenxi Sun also proposed a similar idea, but unlike Ane Blázquez-García, it parameterizes the estimate of the predictive ability of the time point, avoiding repeated training of the model, thereby significantly reducing the computational cost. Although these two methods have shown good results, they both face a common problem: they are based on a unified time point selection strategy, and all features use the same set of time steps. This inflexible strategy cannot select the best data points for each feature individually, which may lead to the loss of important information. This problem is particularly prominent when some features have significant changes at specific time points. For example, in the medical field, when a patient uses human serum albumin and is recorded by a nurse, the data at that moment may play a key role in subsequent clinical decision-making. However, if non-real-time detection indicators (such as blood gas analysis, etc.) are not recorded at this time, then when a unified time point selection strategy is adopted, these key data points may be ignored, affecting the final classification effect of the model.
[0005] In summary, the existing interpolation methods based on irregular time series have the following shortcomings: (1) Traditional methods such as forward interpolation, zero interpolation, and average interpolation ignore the nonlinearity and dynamic changes of data. The interpolation results cannot accurately reflect the real data and are prone to introduce a lot of noise. (2) When deep learning models process multivariate time series, the computational burden of training and inference is particularly heavy, which significantly increases the computational overhead. (3) Many methods adopt a unified time point selection strategy, that is, all features use the same subset of time steps. This approach ignores the change patterns of different features in time series data and cannot select the most appropriate time point for each feature separately. In addition, the unified strategy cannot adapt to the differences between features, especially when some features have sudden changes (such as acute changes in illness). These key moments may be missed, resulting in the interpolation results being unable to accurately reflect the actual changes in the data, affecting subsequent analysis and decision-making. Summary of the invention
[0006] In order to solve the problems of insufficient precision, high computational cost and inflexible data processing methods in existing interpolation methods based on irregular time series, the present invention provides a classification method suitable for irregular time series and its corresponding system and device.
[0007] The present invention is implemented by the following technical solutions:
[0008] A classification method suitable for irregular time series, comprising:
[0009] A classification model is constructed which includes a time-aware interpolator, a front dual self-attention module, an offset selection module, a rear dual self-attention module and a classifier in sequence; the model is used to output corresponding classification results according to the sequence data of the input multi-dimensional features.
[0010] Among them, the time-aware interpolator is used to encode the time input to obtain vectors Q and K, and to encode the observation input to obtain vector V, and then the dot product attention mechanism is combined to complete the data interpolation. The two dual self-attention modules each include a time self-attention layer and a variable self-attention layer for capturing the time relationship and variable relationship, as well as a feedforward layer. The offset selection module is used to generate the offset and the final selected position after the offset according to the initial position and the original feature, and then generate a new feature vector representing the optimized data through the bilinear interpolation method. The classifier is used to decode the input features and generate the probability distribution of different categories through the linear layer and Softmax activation.
[0011] Obtain a large amount of irregular time series with labeled information as sample data, and use the sample data set to train, verify, and test the classification model. Keep the model parameters of the classification model with the best performance and use it to perform the classification task of irregular time series.
[0012] As a further improvement of the present invention, in the time-aware interpolator, the expression of the feature encoding process is:
[0013] ,
[0014] The expression of the data interpolation process is:
[0015] ,
[0016] In the above formula, the time series consists of the observed values of the characteristic parameters and the timestamp information. u The variable category representing the characteristic parameter; represents the eigenvalue vector; Represents the time axis vector in the time series; Indicates time interval information; represents the mask vector used to indicate missing feature values; Indicates the type code; , and Represent the query vector , key vector , value vector The transformation matrix of Represents a matrix operation of element-wise multiplication; represents the output of the time-aware interpolator; Indicates the dimension of the key vector.
[0017] As a further improvement of the present invention, in the temporal self-attention layer of the dual self-attention module, the input features are first layer-normalized and linearly transformed to obtain Q, K, and V vectors. The Q vector and the K vector matrix are multiplied and then processed by the SoftMax layer to obtain a feature vector with a dimension of U×L×L, which is multiplied with the V vector matrix and finally output after element-wise addition operation with the original input.
[0018] As a further improvement of the present invention, in the variable self-attention layer of the dual self-attention module, the input features are first dimensionally changed, and then layer normalization and linear transformation are performed to obtain Q, K, and V vectors; the Q vector and the K vector matrix are multiplied and processed by the SoftMax layer to obtain a feature vector with a dimension of L×U×U, which is multiplied with the V vector matrix, and then element-wise addition is performed with the original input to obtain the result, which is output after dimension transformation.
[0019] As a further improvement of the present invention, the expression of the data processing process of the offset selection module is:
[0020] ,
[0021] In the above formula, subscript i represents the time point of input data H, and subscript j represents the output data time point; Indicates the offset; Indicates the initial selection position of the time series; Indicates the final selected position after the offset operation; offset represents the offset network; represents the output feature of the offset selection module; g represents the bilinear interpolation operation, which satisfies: ; U represents the total number of categories of feature parameters; L Indicates the length of the time series; Represents the mean length of the time dimension in the time series.
[0022] As a further improvement of the present invention, the data processing process of the classifier includes: passing the input features through two additive attention mechanisms to capture the relationship between the feature dimension and the time dimension respectively, and compressing the dimension of the input features from U×L×D to D; the compressed feature vector is input into a linear layer, and finally the probability distribution of each category is obtained through the Softmax activation function.
[0023] As a further improvement of the present invention, the classification model is trained under the guidance of the cosine decay strategy, which is specifically designed to dynamically adjust the learning rate. The loss function of the training phase is as follows:
[0024] ,
[0025] In the above formula, y i is an indicator variable indicating the classification result. If the sample belongs to the category i but y i = 1, otherwise y i = 0; p i The model predicts that the sample belongs to the category i probability; C is the total number of categories.
[0026] As a further improvement of the present invention, the evaluation indicators used by the classification model in the verification stage include AUCROC, AUCPRC and Brier Score.
[0027] The present invention also includes a classification system suitable for irregular time series, which includes a data acquisition module and a data processing module. The data acquisition module is used to obtain the irregular time series to be processed; the data processing module includes a classification model trained in the aforementioned classification method suitable for irregular time series. The classification model is used to generate a corresponding classification result according to the input irregular time series.
[0028] The present invention also includes a classification device suitable for irregular time series, which includes a memory, a processor, and a computer program stored in the memory and running in the processor. When the processor executes the computer program, the classification method suitable for irregular time series as described above is implemented, and then the corresponding classification result is generated according to the input irregular time series.
[0029] The technical solution provided by the present invention has the following beneficial effects:
[0030] The classification model provided by the present invention effectively captures the dynamic relationship between each time point in the time series through the time-aware attention mechanism, significantly improving the quality of data interpolation. Compared with existing interpolation methods, this mechanism has higher computational efficiency and ease of implementation while maintaining high interpolation quality. It is particularly suitable for large-scale data sets and significantly reduces computational overhead.
[0031] The classification model provided by the present invention combines the interrelationships and time dependencies between features through an offset selection strategy, adaptively selecting the most critical data points for each feature, and avoiding the information loss caused by the uniform time point selection of traditional methods. This strategy effectively simplifies the time series representation, reduces redundant data, and thus reduces the complexity and resource consumption of subsequent calculations.
[0032] The present invention combines the interpolation method with the optimization of the classification task, adopts the time-aware interpolator and the offset selection module, and adaptively selects key data points to ensure that the classification model can fully utilize the data features. The accurate data point selection not only reduces the computational burden, but also enhances the generalization ability of the model, reduces the risk of overfitting, and significantly improves the classification performance. This enables the present invention to show excellent robustness and predictive ability in applications with high precision requirements such as medical data containing a large number of irregular time series. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0034] Figure 1 This is a flowchart of the steps of the classification method applicable to irregular time series provided in Example 1 of the present invention.
[0035] Figure 2 This is a network architecture diagram of the classification model constructed in Example 1 of the present invention.
[0036] Figure 3 This is a schematic diagram of the time-aware interpolator in the classification model of Example 1 of the present invention.
[0037] Figure 4 This is a functional module diagram of the dual self-attention module in the classification model of Example 1 of the present invention.
[0038] Figure 5 This is a schematic diagram of the offset selection module in the classification model of Example 1 of the present invention.
[0039] Figure 6 This is a flow chart of the training strategy adopted by the classification model in the training phase in Example 1 of the present invention. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] Example 1
[0042] This embodiment provides a classification method suitable for irregular time series. This method introduces a time-aware attention mechanism into the traditional classification model to capture the relationship between each time point in the time series, thereby guiding the data interpolation work to reduce the amount of calculation. In addition, this embodiment also introduces a data point selection strategy based on data interpolation, and adaptively selects the most important data points according to the attributes and time dependencies of different features to ensure that key data points are retained, thereby avoiding information loss and improving classification accuracy. The method provided in this embodiment combines interpolation with classification task optimization, and adopts time-aware interpolation and offset selection strategies to adaptively select key data points, ensuring that the classification model can fully utilize data features, thereby significantly improving classification performance. Specifically, if Figure 1 As shown, the classification method applicable to irregular time series provided in this embodiment includes:
[0043] 1. Construction of classification model
[0044] like Figure 2 As shown, a classification model including a time-aware interpolator, a front dual self-attention module, an offset selection module, a rear dual self-attention module and a classifier is constructed in sequence; it is used to output the corresponding classification result according to the sequence data of the input multi-dimensional features. In the classification model, the original sequence data is first input into the time interpolator, and the time-aware interpolator is used to encode the time input to obtain vectors Q and K, and encode the observation input to obtain vector V, and then the dot product attention mechanism is combined to complete the interpolation of the missing data to obtain a complete feature representation. The feature information after data interpolation is input into the front dual self-attention module to capture the time relationship and variable relationship in the feature information, and generate a feature representation containing global features and time information. Next, the output of the front dual self-attention module enters the offset selection module, and the offset selection module is used to generate an offset and a final selected position after the offset according to the initial position and the original feature, and then a new feature vector representing the optimized data is generated by the bilinear interpolation method. The feature information processed by the offset selection module selects data points that can better express the sequence information and simplify the sequence representation. The output of the offset selection module enters the post-double self-attention module, which continues to extract the time relationship and variable relationship in the feature vector optimized by data offset and bilinear interpolation method; and outputs the processing results to the classifier. The classifier in this embodiment is used to decode the input features and generate probability distributions of different categories through linear layers and Softmax activation; and then obtain the final classification result. In order to make the scheme details of the classification model designed in this embodiment clearer, the following describes in detail each key data processing model contained in the classification model:
[0045] 1.1 Time-aware interpolator
[0046] In the solution provided in this embodiment, the working principle of the time-aware interpolator is as follows: Figure 3 As shown. Specifically, the time-aware interpolator contains a feature encoding unit and an attention module based on the dot product attention mechanism. The original input received by the time-aware interpolator is a multidimensional time series. The multidimensional time series s is a discrete data consisting of the observed values of multiple feature parameters and their corresponding timestamp information. The basic format of the multidimensional time series s is as follows:
[0047] ,
[0048] In the above formula, Indicates that the u-th characteristic parameter is i The observed value at time, t i Indicates i Timestamp information of a moment; L Indicates the length of a multidimensional time series; U Indicates the total number of categories of feature parameters.
[0049] In the time-aware interpolator, the feature encoding unit first encodes the time input T in the input multidimensional time series s to obtain the query vector Q and the key vector K; and encodes the observation value input (also known as the feature value vector) X to obtain the value vector V. Among them, the data format of the time input T contained in the multidimensional time series s is as follows:
[0050] .
[0051] The data format of the observation input X contained in the multidimensional time series s is as follows:
[0052] .
[0053] In addition, in the time-aware interpolator of this embodiment, a mask vector M is also used to characterize the situation where there are missing values in the eigenvalue vector X. The data format of the mask vector M is as follows:
[0054] ,
[0055] In the above formula, each element in the mask vector M The value of is determined by the following formula:
[0056] ;
[0057] Specifically, the expression of the feature encoding process is:
[0058] ,
[0059] In the above formula, Indicates u Eigenvalue vector of class features; Indicates the time series u The time axis vector of the class features; Indicates u Time interval information of class features; Indicates the u Mask vector for missing class eigenvalues; Indicates u Class feature type encoding; , and Respectively represent u Query vector of class features , key vector , value vector The transformation matrix of .
[0060] The time interval information Δ is used to represent the time interval between the current time and the previous observation value. u , the corresponding time interval information Obtained by the following formula:
[0061] ,
[0062] The type code E is calculated by the following formula:
[0063]
[0064] In the time-aware interpolator of this embodiment, the Q, K, and V vectors generated by the encoding unit are combined; the dot product attention mechanism first performs a dot product operation on Q and K to obtain an L×L attention matrix, and uses a mask vector M to mask the part without observation values in the second dimension, so that the attention matrix becomes a relationship matrix between all observation time points and time points with observation values. Next, this relationship matrix is used to supplement the missing data in V to obtain a complete feature representation, thereby completing the data interpolation task.
[0065] Specifically, the expression of the data interpolation process is:
[0066]
[0067] In the above formula, Represents a matrix operation of element-wise multiplication; represents the output of the time-aware interpolator; Indicates the dimension of the key vector.
[0068] 1.2. Dual Self-Attention Module
[0069] The classification model provided in this embodiment includes two identical dual self-attention modules, one of which is located after the time-aware interpolator and the other is located after the offset selection module. The two dual self-attention modules are used to extract the time relationship and feature relationship contained in the complete feature representation after data interpolation by the time-aware interpolator, and the time relationship and feature relationship contained in the feature vector after data offset selection and optimization by the offset selection module.
[0070] like Figure 4 As shown in the figure, the dual self-attention module includes a temporal self-attention layer, a variable self-attention layer, and a feedforward layer. The temporal self-attention layer is used to capture the temporal relationship in feature representation; the variable self-attention layer is used to capture the variable relationship in feature representation. The feedforward layer introduces necessary nonlinearity to the model, increases the depth of the model, and enables the model to learn complex function mappings; it helps to improve the model's expressiveness and generalization capabilities.
[0071] In detail, in the temporal self-attention layer of the dual self-attention module provided in this embodiment, the input features are first layer-normalized and linearly transformed to obtain Q, K, and V vectors, and the Q vector and K vector matrices are multiplied and processed by the SoftMax layer to obtain a feature vector with a dimension of U×L×L, which is multiplied by the V vector matrix, and finally output after element addition operation with the original input. In the variable self-attention layer of the dual self-attention module, the input features are first dimensionally changed, and then layer-normalized and linearly transformed to obtain Q, K, and V vectors; the Q vector and K vector matrices are multiplied and processed by the SoftMax layer to obtain a feature vector with a dimension of L×U×U, which is multiplied by the V vector matrix, and then element addition operation is performed with the original input, and the result is output after dimension transformation.
[0072] 1.3. Offset selection module
[0073] In this embodiment, the dimension of the original input of the offset selection module is U×L×D, where D is the dimension of the feature information contained in each feature parameter. Figure 5 As shown, during the data processing of the offset selection module, for each feature, an initial selection position is first generated according to the sequence length , the window interval is 2 and the length is the median of the sequence length.
[0074] The data at these initially selected locations are then fed into an offset network, which identifies data points with poor interpolation quality based on the global temporal information in the data and generates corresponding offsets ; Add the offset to the index of the initial selection window to get the final selection position.
[0075] Specifically, the expression for generating the offset is as follows:
[0076] ,
[0077] Among them, the subscript j represents the time point of output data; Indicates the initial selection position of the time series; The received input features of the corresponding offset selection module; offset Represents an offset network.
[0078] The expression that determines the final selected position is as follows:
[0079] ,
[0080] In the above formula, Indicates the final selected position after the offset operation.
[0081] Finally, the value corresponding to the final window is calculated by linear interpolation to complete the offset selection of the data point. In this embodiment, the offset selection of the data point is completed by using a bilinear interpolation method, and the corresponding expression is as follows:
[0082] ,
[0083] In the above formula, the subscript i Indicates the time point of input data; represents the output features of the offset selection module; represents the mean length of the time dimension in the time series; g represents the bilinear interpolation operation, which satisfies the following formula:
[0084] .
[0085] After the above processing, the offset selection module can discard data points with poor interpolation quality while ensuring the consistency of the output results.
[0086] 1.4 Classifier
[0087] The classifier provided in this embodiment includes a decoding unit composed of an additive attention mechanism; the dimension of the output of the post-double self-attention module received by the classifier is U×L×D; in the decoding unit, the input of the classifier passes through two additive attention mechanisms to capture the relationship between the feature dimension and the time dimension respectively. After each additive attention operation, the dimension of the data is gradually reduced, and important temporal features or spatial features are extracted. After two additive attention processes, the dimension of the input is compressed from U×L×D to D, that is, the final representation of each sample is a D-dimensional vector.
[0088] Next, the D-dimensional vector output by the decoding unit is input into a linear layer. The input dimension of the linear layer is D, and the output dimension is the number of categories. The linear layer finally obtains the probability distribution of each category through the Softmax activation function, thereby completing the classification task.
[0089] 2. Training and Application of Classification Model
[0090] A large amount of irregular time series with labeled information is obtained as sample data, and the sample data set is divided into training set, validation set and test set, which are used to train, validate and test the classification model. The model parameters of the classification model with the best performance are retained and used to perform the classification task of irregular time series.
[0091] In this embodiment, the basic format of the sample data is (s, y), where s represents a multidimensional time series, and y represents a label corresponding to the multidimensional time series. This embodiment collects a large amount of sample data and preprocesses the data set to remove outliers and invalid values contained in the time series. Next, the collected sample data set is divided. This embodiment divides the sample data set into a training set, a validation set, and a test set in a ratio of 64:16:20. Among them, the training set is used to train the classification model, the validation set is used to select the best model configuration, and the test set is used to calculate the performance indicators of the classification model.
[0092] In the solution of this embodiment, the training strategy adopted by the classification model in the training phase is as follows:
[0093] The overall model is trained under the guidance of the cosine decay strategy, with the initial learning rate set to 0.001, and the Adam optimizer is used for parameter update. During each round of training, the model is verified in real time using the validation set to monitor the performance changes of the model. Specifically, during the training process, if the performance on the validation set does not improve significantly within 10 consecutive rounds (i.e., 10 training cycles), the training will be stopped in advance to prevent overfitting and save computing resources. In addition, in order to improve the generalization ability of the model, appropriate regularization strategies (such as Dropout or L2 regularization) are also used in the training process to further avoid overfitting. The cosine decay learning rate scheduler gradually reduces the learning rate as the training progresses, thereby helping the model to better converge to the optimal solution.
[0094] Specifically, the loss function of the solution in this embodiment during the training phase is as follows:
[0095] ,
[0096] In the above formula, y i is an indicator variable indicating the classification result. If the sample belongs to the category ibut y i = 1, otherwise y i = 0; p i The model predicts that the sample belongs to the category i probability; C is the total number of categories.
[0097] As a further improvement of the present invention, the evaluation indicators used by the classification model in the verification stage include AUCROC, AUCPRC and Brier Score.
[0098] The present invention also includes a classification system suitable for irregular time series, which includes a data acquisition module and a data processing module. The data acquisition module is used to obtain the irregular time series to be processed; the data processing module includes a classification model trained in the aforementioned classification method suitable for irregular time series. The classification model is used to generate a corresponding classification result according to the input irregular time series.
[0099] The present invention also includes a classification device suitable for irregular time series, which includes a memory, a processor, and a computer program stored in the memory and running in the processor. When the processor executes the computer program, the classification method suitable for irregular time series as described above is implemented, and then the corresponding classification result is generated according to the input irregular time series.
[0100] During the training process, to ensure that the best model is selected, the performance of the current model on the validation set is evaluated after each round of training and compared with the previous best performance. If the current model performs better, the best model is updated. When the training process triggers early stopping, the model will be restored to the state with the best performance on the validation set, ensuring that the best model is finally used. After training is completed, the best performing model on the validation set is saved and used in the subsequent testing phase for evaluation to ensure the reliability and accuracy of the test results. The specific process is as follows Figure 6 As shown,
[0101] In the test phase, it is only necessary to load the weight parameter file of the trained optimal model into the model. Then input the test set into the model and wait for the output result. The evaluation indicators used in the test phase of this embodiment include AUC-ROC, AUC-PRC and Brier Score. Among them, the indicator AUC-ROC is used to determine whether the model can effectively distinguish the target event from the background information in the time series under different thresholds, and measure the model's ability to distinguish between positive and negative classes. The indicator AUC-PRC is used to determine the balance between the precision and recall rate of the model when processing positive events in the time series, especially in the case of class imbalance, to evaluate the model's ability to effectively identify and accurately predict a few target events. The indicator Brier Score is used to evaluate the probability prediction accuracy of the model in time series prediction, measure the accuracy and reliability of the model in estimating the probability of event occurrence, and reflect whether the model can accurately predict the probability of the target event and provide the correct confidence.
[0102] Specifically, AUC-ROC is an indicator for evaluating the performance of a classification model by calculating the area under the ROC curve. AUC-ROC is a commonly used indicator, especially in the case of class imbalance. The ROC curve plots the relationship between the true positive rate (TPR) and the false positive rate (FPR) by continuously changing the threshold; the calculation formulas for TPR and FPR are as follows:
[0103] ,
[0104] In the above formula, TP represents true positive samples, FN represents false negative samples, FP represents false positive samples, and TN represents true negative samples.
[0105] AUC-PRC is an indicator for evaluating the performance of a classification model by calculating the area under the precision-recall curve. In the case of class imbalance, AUC-PRC is a commonly used indicator for evaluating models. The precision-recall curve plots the relationship between precision and recall by continuously changing the threshold; the calculation formulas for Precision and Recall are as follows:
[0106] ,
[0107] Brier Score is an indicator used to evaluate the accuracy of probability prediction, which is applicable to regression or classification problems with probability output corresponding to the present invention. This indicator measures the prediction quality of the model by calculating the mean square error between the true label and the predicted probability. The calculation formula is as follows:
[0108]
[0109] In the above formula,N Indicates the number of samples; V Indicates the total number of categories; Indicates i The samples belong to j The predicted probability of each category; Indicates i The samples belong to j The actual result of a category is 1 if the sample belongs to that category, otherwise it is 0.
[0110] In the actual application process of the scheme of this embodiment, the method can be designed as a computer software system, that is, a classification system suitable for irregular time series, which includes a data acquisition module and a data processing module. The data acquisition module is used to obtain the irregular time series to be processed; the data processing module includes a classification model trained in the aforementioned classification method suitable for irregular time series. The classification model is used to generate a corresponding classification result based on the input irregular time series.
[0111] The typical application scenario of this embodiment is the medical field. For example, this method can be used to analyze patient data and assist doctors in real-time assessment of patient risks. Since this embodiment can be applied to generate corresponding classification results in combination with multi-dimensional, irregular time series data, it can be used to analyze complex time series data composed of multiple factors such as age, medical history, and physiological indicators, and then use the data interpolation and optimization strategy of the present invention in the deep learning algorithm to improve the accuracy of classification prediction. In such scenarios, the rapid assessment tool provided by this embodiment improves the quality of medical decision-making, helps optimize the allocation of medical resources, and ensures that high-risk patients receive timely intervention.
[0112] Example 2
[0113] On the basis of Example 1, this embodiment further provides a classification device applicable to irregular time series. It includes a memory, a processor, and a computer program stored in the memory and running in the processor. When the processor executes the computer program, the classification method applicable to irregular time series as described above is implemented, and then the corresponding classification result is generated according to the input irregular time series. The classification device provided in this embodiment is an actual product solution for implementing the method in Example 1.
[0114] Specifically, the classification device in this embodiment is essentially a computer device, which can be an embedded device integrated into a medical monitoring instrument to generate classification results in combination with real-time monitoring data. An independent device independent of the medical monitoring instrument can also be used. For example, the computer device in this embodiment can be an intelligent terminal capable of executing a program, a tablet computer, a laptop computer, a desktop computer, a rack server, a blade server, a tower server or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc.
[0115] The computer device indicated in this embodiment includes at least but is not limited to: a memory and a processor that can be connected to each other through a system bus. Among them, the memory (i.e., a readable storage medium) includes a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory can be an internal storage unit of a computer device, such as a hard disk or a memory of the computer device. In other embodiments, the memory can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the memory is generally used to store an operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various types of data that have been output or are to be output.
[0116] The processor may be a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor is generally used to control the overall operation of a computer device. In this embodiment, the processor is used to run program codes stored in a memory or process data.
[0117] Verification experiment
[0118] In order to verify the performance of the classification method for irregular time series provided by the present invention, the technicians formulated an experimental plan to verify the scheme on two data sets, PhysioNet Challenge 2012 and MIMIC-IV, and selected multiple existing classification models as the control group to compare the prediction performance with this case.
[0119] 1. Dataset
[0120] The PhysioNet Challenge 2012 dataset collects 4 types of demographic information and 37 types of physiological signals collected from patients in the first 48 hours of their stay in the ICU. The task is to predict in-hospital mortality. The dataset has three subsets, each containing 4,000 instances. To ensure the scientificity and fairness of the experiment, this experiment refers to previous related experiments and uses different subset combinations to construct three datasets, named Physio4000, Physio8000, and Physio12000.
[0121] The MIMIC-IV dataset is a health data record of patients in the Intensive Care Unit (ICU) of Beth Israel Deaconess Medical Center from 2008 to 2019. This experiment extracted 6636 instances from it, containing 49 physiological signal information collected within 24 hours. The task of this dataset is to predict whether the patient will die in the next 24 hours based on the data within the 24-hour time window.
[0122] 2. Control group plan
[0123] (1) RNN is a classic model in time series analysis. Since it relies on non-missing inputs at uniform time intervals, we use hourly aggregation and data interpolation (forward filling, mean interpolation), and thus divide it into RNN-Mean and RNN-Forward.
[0124] (2) GRU, as a variant of RNN, uses the same settings as RNN and is divided into GRU-Mean and GRU-Forward.
[0125] (3) Warpformer performs multi-scale modeling of irregular clinical sequences and emphasizes the irregularity within the sequence and the differences between sequences.
[0126] (4) MTSFormer embeds data from multiple perspectives and adaptively learns the irregular dynamics of multivariate time series.
[0127] (5) DNA-T focuses on enhancing the learning of local information and capturing the correlation of local changes in patient status.
[0128] 3. Experimental results and analysis
[0129] In this experiment, three indicators, AUC-ROC, AUC-PRC and Brier Score, were selected to evaluate the performance of the scheme of the present invention and the control group scheme on various data sets. The experimental results are shown in the following table:
[0130] Table 1: Performance of the present invention and control group solutions
[0131]
[0132] By analyzing the experimental data in the above table, it can be found that the three indicators of the solution of the present invention have achieved almost the best results on the four data sets, which reflects the outstanding performance of the present invention in the classification task of irregular time series.
[0133] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A classification method suitable for irregular time series, characterized in that: It includes: Construct a classification model including a time-aware interpolator, a front dual self-attention module, an offset selection module, a rear dual self-attention module and a classifier in sequence; It is used to output the corresponding classification results according to the sequence data of the input multi-dimensional features; The time-aware interpolator is used to encode the time input to obtain vectors Q and K, and to encode the observation input to obtain vector V, and then combine the dot product attention mechanism to complete data interpolation; the two dual self-attention modules include a time self-attention layer and a variable self-attention layer for capturing time and variable relationships, and a feedforward layer; the offset selection module is used to generate an offset and a final selected position after the offset according to the initial position and the original feature, and then generate a new feature vector representing the optimized data through a bilinear interpolation method; the classifier is used to decode the input features and generate probability distributions of different categories through a linear layer and Softmax activation; in the time-aware interpolator, the expression of feature encoding is: , The expression for data interpolation is: , In the above formula, the time series consists of the observed values of the characteristic parameters and the timestamp information. u The variable category representing the characteristic parameter; represents the eigenvalue vector; Represents the time axis vector in the time series; Indicates time interval information; represents the mask vector used to indicate missing feature values; Indicates the type code; , and Represent the query vector , key vector , value vector The transformation matrix of Represents a matrix operation of element-wise multiplication; represents the output of the time-aware interpolator; represents the dimension of the key vector; A large amount of irregular time series with labeled information is obtained as sample data, and the classification model is trained, verified and tested using the sample data set. The model parameters of the classification model with the best performance are retained and used to perform the classification task of irregular time series.
2. The classification method for irregular time series according to claim 1, characterized in that: In the temporal self-attention layer of the dual self-attention module, the input features are first layer-normalized and linearly transformed to obtain Q, K, and V vectors. The Q vector and the K vector matrix are multiplied and processed by the SoftMax layer to obtain a feature vector with a dimension of U×L×L, which is multiplied with the V vector matrix and finally output after element-wise addition operation with the original input.
3. The classification method applicable to irregular time series as claimed in claim 1, characterized in that: In the variable self-attention layer of the dual self-attention module, the input features are first dimensionally changed, and then layer normalization and linear transformation are performed to obtain Q, K, and V vectors; the Q vector and the K vector matrix are multiplied and processed by the SoftMax layer to obtain a feature vector with a dimension of L×U×U, which is multiplied with the V vector matrix, and then element-wise addition is performed with the original input to obtain the result and output it after dimension transformation.
4. The classification method for irregular time series as claimed in claim 1, characterized in that: The expression of the data processing process of the offset selection module is: , In the above formula, the subscript i Indicates the time point of input data, subscript j Indicates the time point of output data; Indicates the offset; Indicates the initial selection position of the time series; Indicates the final selected position of the time series after the offset operation; offset represents the offset network; represents the output feature of the offset selection module; g represents the bilinear interpolation operation, which satisfies: ; U represents the total number of categories of feature parameters; L Indicates the length of the time series; Represents the mean length of the time dimension in the time series.
5. The classification method applicable to irregular time series as claimed in claim 4, characterized in that: The data processing process of the classifier includes: passing the input features through two additive attention mechanisms to capture the relationship between the feature dimension and the time dimension respectively, and compressing the dimension of the input features from U×L×D to D; the compressed feature vector is input into a linear layer, and finally the probability distribution of each category is obtained through the Softmax activation function.
6. The classification method for irregular time series as claimed in claim 1, characterized in that: The classification model is trained under the guidance of the cosine decay strategy, and the loss function of the training phase is as follows: , In the above formula, y i is an indicator variable indicating the classification result. If the sample belongs to the category i but y i = 1, otherwise y i = 0; p i The model predicts that the sample belongs to the category i probability; C is the total number of categories.
7. The classification method applicable to irregular time series as claimed in claim 1, characterized in that: The evaluation indicators used by the classification model in the testing phase include AUCROC, AUCPRC and Brier Score.
8. A classification system suitable for irregular time series, characterized in that: It includes a data acquisition module and a data processing module; the data acquisition module is used to obtain the irregular time series to be processed; the data processing module includes a classification model that has been trained in the classification method applicable to irregular time series as described in any one of claims 1-7; the classification model is used to generate corresponding classification results based on the input irregular time series.
9. A classification device for irregular time series, comprising a memory, a processor, and a computer program stored in the memory and running in the processor, characterized in that: When the processor executes the computer program, it implements the classification method applicable to irregular time series as described in any one of claims 1 to 7, and further generates a corresponding classification result according to the input irregular time series.
Citation Information
Patent Citations
Multivariate time series classification method and device based on bidirectional self-attention
CN115563537A
Charging pile fault detection method based on probability sparse self-attention mechanism
CN118821031A