Time sequence similarity evaluation method and device based on DTW and K nearest neighbor algorithm

By combining DTW algorithm with deep learning, dynamic and learnable DTW alignment paths, and combining K nearest neighbor algorithms, the shortcomings of the existing technology in processing complex and nonlinear time series data are solved, and efficient and accurate time series classification and prediction are achieved.

CN120046027APending Publication Date: 2025-05-27XIAMEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510208577.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing time series analysis methods are insufficient in processing complex, nonlinear data, especially in capturing dynamic changes and long-term trends of data, and the application of deep learning models in large-scale data sets is limited by the problems of high computational costs and long training cycles.

Method used

Combining dynamic time regularization (DTW) algorithm and deep learning, DTWNet makes the alignment path of DTW dynamic and learnable, and combining K nearest neighbor algorithm for time series similarity evaluation and classification.

Benefits of technology

It improves computational efficiency and classification accuracy, adapts to different types of time series data, overcomes the adaptability problem of traditional DTW in nonlinear complex data, and provides good interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046027A_ABST
    Figure CN120046027A_ABST
Patent Text Reader

Abstract

The invention discloses a DTW and K nearest neighbor algorithm-based time sequence similarity evaluation method and device, and relates to the field of time sequence analysis, and the method comprises the steps: S1, collecting target time sequence data in real time, and carrying out the preprocessing of the data; s2, constructing a deep learning model comprising a DTW Function module, a DTW Layer and an optimizer; s3, inputting the preprocessed data and the collected historical data into a deep learning model to obtain the similarity of time sequences; and S4, based on the similarity of the time series, classifying the preprocessed data by using a K nearest neighbor algorithm, and completing time series similarity evaluation. According to the method, the nonlinear alignment capability of the DTW is combined with the learnable DTW kernel in the deep learning model, so that the calculation efficiency and classification precision of the large-scale time sequence data are obviously improved, and meanwhile, a transparent and interpretable classification result is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of time series analysis, and particularly to a method and apparatus for evaluating the similarity of time series based on the DTW and K-nearest neighbor algorithms. Background Art

[0002] Time series analysis is the core content of modern metrological statistical analysis and has wide applications in data analysis in fields such as finance, healthcare, and industrial control. With the rapid development of big data and machine learning technologies, time series analysis still faces many challenges when dealing with large-scale and complex non-linear data, especially showing limitations in capturing non-linear patterns and long-term trends in the data.

[0003] Traditional time series analysis methods, such as the autoregressive integrated moving average model (ARIMA) and seasonal decomposition, are usually based on linear assumptions. However, actual time series data often exhibits complex non-linear characteristics, especially in complex scenarios such as financial market fluctuations, biological signal detection, and climate change monitoring. These methods are insufficient in accurately capturing the dynamic changes of the data. At the same time, due to relying on analysis windows of fixed sizes, they often ignore the long-term trends and periodic changes in the data and are difficult to comprehensively reveal the overall characteristics and internal laws of the data.

[0004] Current machine learning methods face many challenges when dealing with complex time series data. For example, the support vector machine (SVM) performs excellently when dealing with linearly separable data sets, but performs poorly for highly complex non-linear time series, especially prone to overfitting in a high-dimensional data environment. The recurrent neural network (RNN) can retain the temporal information of the time series, but is limited by the problems of gradient vanishing and gradient explosion and has limited performance when dealing with long time series. Although the long short-term memory network (LSTM) improves some deficiencies by introducing a gating mechanism, when dealing with ultra-long time series, it still faces problems such as huge consumption of computing resources and too long training cycles. In addition, deep learning models such as LSTM and gated recurrent unit (GRU) perform excellently in complex time series prediction, but the high computing cost and long training cycle limit their practical applications in large-scale data sets. At the same time, these models are regarded as "black box operations", and their internal decision-making mechanisms are difficult to explain, which brings significant challenges in fields such as medical diagnosis that require high transparency and interpretability.

[0005] The Dynamic Time Warping (DTW) algorithm is a tool for measuring the similarity of time series. By means of a "one-to-many" matching method, it overcomes the limitation of the Euclidean distance that can only match point to point. DTW can adapt to the displacement and amplitude changes in time series and has strong robustness to asynchrony. However, when dealing with large-scale data, the computational complexity of DTW is relatively high, which becomes the main bottleneck for its wide application. In addition, DTW may ignore local morphological features when aligning sequences, resulting in insufficient accuracy when dealing with complex fluctuation patterns.

[0006] To address the above problems, the present invention proposes a time series cross-linking method combining the DTW algorithm and a neural network (DTWNet). It is intended to combine the non-linear alignment ability of DTW with the learnable DTW kernel in DTWNet to process large-scale and complex time series data, thereby significantly improving the computational efficiency and classification accuracy. Summary of the Invention

[0007] To address the above problems, the present invention proposes a time series similarity evaluation method and device based on DTW and the K-nearest neighbor algorithm. Utilizing the DTW algorithm in combination with deep learning, it is specifically designed for complex and non-linear time series data. For problems such as inconsistent sequence lengths, asynchrony, and non-linear changes, it provides an efficient and accurate solution. By leveraging the non-linear alignment function of the DTW algorithm, the present invention effectively overcomes the deficiencies of the Euclidean distance in dealing with asynchronous sequences. Combining stream feature capture and the design of a differentiable kernel in deep learning, it achieves efficient time series classification and prediction, thus making up for the limitations of the prior art in dealing with non-linear data or long time series.

[0008] On the one hand, the time series similarity evaluation method based on DTW and the K-nearest neighbor algorithm is as follows:

[0009] S1. Real-time collect the target time series data and preprocess the data.

[0010] S2. Construct a deep learning model including a DTW Function module, a DTW Layer, and an optimizer; the DTWFunction module is a differentiable neural network component based on the DTW algorithm, which realizes the forward path calculation and the backward gradient calculation; the DTW Layer encapsulates the DTW algorithm into a neural network layer and extracts time series features layer by layer; the optimizer updates the parameters of the deep learning model.

[0011] S3. Input the preprocessed data and the collected historical data into the deep learning model to obtain the similarity of the time series.

[0012] S4. Classify the preprocessed data using the K-nearest neighbor algorithm based on the similarity of time series to complete the time series similarity evaluation.

[0013] Preferably, the preprocessing includes data standardization and normalization, noise processing, and missing value filling; the noise processing uses the moving average method; the missing value filling uses linear interpolation and automatic filling techniques.

[0014] Preferably, the optimizer is the AdamW optimizer.

[0015] Preferably, the specific calculation of the reverse gradient is as follows:

[0016] Determine an optimal alignment path as a fixed path, and then calculate the reverse gradient only on this fixed path; the reverse gradient calculation formula is expressed as:

[0017]

[0018] where x represents the source time series; y represents the time series to be aligned; represents the partial derivative of x; dtw 2 represents the square of the dtw function.

[0019] Preferably, the objective function of the deep learning model is:

[0020] Loss = C(i,j) + λR(C)

[0021] where Loss represents the objective function; C(i,j) represents the minimum cumulative distance from point (0,0) to position (i,j); λ represents the regularization parameter; R(C) represents the smoothing regularization term.

[0022] Preferably, the recurrence formula for the minimum cumulative distance is expressed as:

[0023] C(i,j) = D(i,j) + min(C(i - 1,j), C(i,j - 1), C(i - 1,j - 1))

[0024] where D(i,j) is a trainable distance matrix representing the distance metric between sequence A and sequence B at position (i,j); min(·) represents taking the minimum value; C(i - 1,j) represents the cumulative cost in the vertical direction; C(i,j - 1) represents the cumulative cost in the horizontal direction; C(i - 1,j - 1) represents the cumulative cost in the diagonal direction.

[0025] Preferably, the trainable distance matrix is expressed as:

[0026]

[0027] Among them, f(·) and h(·) respectively represent the feature transformation functions learned through deep neural networks; A i represents the i-th point of sequence A; B j represents the j-th point of sequence B; ||·|| represents the Euclidean distance function.

[0028] Preferably, the deep learning model uses a multi-scale approximation method for FastDTW optimization.

[0029] Preferably, for the time series similarity, the K-nearest neighbor algorithm is used to classify the preprocessed data, including the following steps:

[0030] S41, based on the time series similarity, select the k historical data that are most similar to the time series of the preprocessed data as the nearest neighbor samples;

[0031] S42, perform weight assignment on the k nearest neighbor samples, and the weight calculation formula is as follows:

[0032]

[0033] Among them, w α represents the weight of the α-th nearest neighbor sample; dist(A, B α ) represents the time series similarity between the current sequence A and the nearest neighbor sample B α ;

[0034] S42, perform weighted voting on the k nearest neighbor samples to obtain the predicted category of the preprocessed data; the formula is as follows:

[0035]

[0036] Among them, c represents the sequence category; arg max represents the independent variable when the function obtains the maximum value; K represents the set of k nearest neighbor samples; represents the predicted category; w m represents the weight of the m-th nearest neighbor sample; y m represents the category label of the m-th nearest neighbor sample; Π represents the indicator function, which returns 1 when the condition is satisfied and 0 otherwise.

[0037] On the other hand, a time series similarity evaluation device based on DTW and the K-nearest neighbor algorithm includes the following:

[0038] A data acquisition and preprocessing module, which is used to collect target time series data in real time and preprocess the data;

[0039] A model construction module for constructing a deep learning model including a DTW Function module, a DTW Layer, and an optimizer; the DTW Function module is a differentiable neural network component based on the DTW algorithm, which realizes forward path calculation and reverse gradient calculation; the DTW Layer encapsulates the DTW algorithm into a neural network layer and extracts time series features layer by layer; the optimizer updates the parameters of the deep learning model;

[0040] A time series similarity acquisition module for inputting the preprocessed data and the collected historical data into the deep learning model to obtain the similarity of the time series;

[0041] A K-nearest neighbor algorithm category prediction module for classifying the preprocessed data using the K-nearest neighbor algorithm based on the similarity of the time series to complete the evaluation of the time series similarity.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] (1) In the present invention, through deep learning, the alignment path of DTW is made dynamic and learnable, so as to adapt to different types of time series data, and the adaptability problem of traditional DTW in non-linear complex data is solved;

[0044] (2) The present invention combines deep learning feature extraction and the KNN classification algorithm, and uses a similarity weighting mechanism to achieve efficient classification of complex time series;

[0045] (3) The present invention combines the fast DTW algorithm and the multi-scale feature extraction of the deep learning model, which improves the calculation efficiency while ensuring the application performance of the model on large-scale data sets;

[0046] (4) The classification result of the present invention is not only a single output, but also provides a similarity comparison with historical data, has good interpretability, is convenient for users to refer to and verify in detection, and even adjusts the result in combination with manual experience;

[0047] (5) The classification and prediction accuracy of the present invention will continuously improve with the accumulation of historical data, and higher accuracy and reliability can be achieved by continuously optimizing the database. Description of the Drawings

[0048] The following further describes the present invention in detail with reference to the drawings;

[0049] Figure 1 It is a flowchart of the time series similarity evaluation method based on the DTW and K-nearest neighbor algorithms according to the embodiment of the present invention;

[0050] Figure 2Algorithm flowchart of the time series similarity evaluation method based on DTW and K-nearest neighbor algorithms according to the embodiments of the present invention;

[0051] Figure 3 Flowchart of the DTW algorithm combined with deep learning of the time series similarity evaluation method based on DTW and K-nearest neighbor algorithms according to the embodiments of the present invention;

[0052] Figure 4 Schematic diagram of the DTW algorithm of the time series similarity evaluation method based on DTW and K-nearest neighbor algorithms according to the embodiments of the present invention;

[0053] Figure 5 Original data graph of the bacterial pressure time series in the first embodiment of the present invention; Each graph represents the pressure time series image of a kind of bacteria in eight states, and there are eight kinds of bacteria in total;

[0054] Figure 6 Data graph of the bacterial pressure time series before algorithm processing in the first embodiment of the present invention;

[0055] Figure 7 Data graph of the bacterial pressure time series after algorithm processing in the first embodiment of the present invention;

[0056] Figure 8 Accuracy comparison graph of the bacterial detection experiment algorithm in the first embodiment of the present invention;

[0057] Figure 9 Running time comparison graph of the bacterial detection experiment algorithm in the first embodiment of the present invention;

[0058] Figure 10 Data graph of the temperature time series in a certain area in the second embodiment of the present invention;

[0059] Figure 11 Accuracy comparison graph of the climate prediction algorithm in the second embodiment of the present invention;

[0060] Figure 12 Structure block diagram of the time series similarity evaluation device based on DTW and K-nearest neighbor algorithms according to the embodiments of the present invention. Detailed implementation manners

[0061] The present invention will be further described below through specific implementation manners.

[0062] See Figure 1 and Figure 2 shown, the time series similarity evaluation method based on DTW and K-nearest neighbor algorithms.

[0063] S1, Real-time collect the target time series data and preprocess the data.

[0064] In bacterial detection, the target time series data is the time series of the surface mechanical expression (or direct surface tension) of bacteria; in climate prediction, the target time series data is the time-temperature series; in financial market analysis, the target time series data is the time series of stock price fluctuations.

[0065] The preprocessing steps are as follows:

[0066] S11, Data standardization and normalization: Standardize and normalize the time series data to make the data have zero mean and unit variance. The standardization formula is as follows:

[0067]

[0068] Among them, X is the original data, μ is the mean, σ is the standard deviation, and X′ is the standardized data. This processing method can reduce the dimensionality difference between features, thereby improving the convergence efficiency of the deep learning model during training.

[0069] S12, Noise processing: Use the moving average method to smooth and denoise the data to reduce the interference of noise on the model performance. These methods can help the model capture key sequence features more effectively during the feature learning process. It should be noted that other noise processing techniques can also be used, including Exponential Weighted Moving Average (EWMA) and Kalman filtering, which are specifically set according to needs and are not restricted in this method.

[0070] S13, Missing value filling: Use linear interpolation and automatic filling techniques in deep learning to process missing values to ensure the integrity and continuity of the time series data. These methods effectively prevent the adverse effects caused by data missing on model training and inference.

[0071] S2, Build a deep learning model based on the DTW algorithm.

[0072] The specific application process of the dynamic time warping (DTW) algorithm combined with deep learning is as Figure 3 shown. The DTW process combined with deep learning includes three main parts: 1) The DTW Function module implements forward path calculation and reverse gradient optimization; 2) The DTW Layer extracts time series features layer by layer; 3) With the support of the AdamW optimizer, optimize the model parameters through an adaptive learning rate, and finally achieve efficient path alignment and classification. Among them, the basic principle diagram of the DTW algorithm is shown in Figure 4 shown.

[0073] Specifically, the DTW Function is based on the DTW algorithm and inherits from the Function class of PyTorch, making DTW differentiable and transforming it into a differentiable neural network component. The DTW Layer is the network layer part of the model, encapsulating the DTW algorithm into a neural network layer, responsible for feature extraction and transformation. AdamW is the optimizer, responsible for parameter update.

[0074] The deep learning module introduced in this method has the following improvements:

[0075] (1) Construct a trainable distance matrix

[0076] The traditional DTW algorithm calculates the similarity between each point in the time series based on the Euclidean distance, and its calculation formula is

[0077] D(i,j) = |A i - B j |

[0078] where A i and B j are the i-th and j-th points of the two sequences respectively.

[0079] This method combines deep learning techniques to design a trainable distance matrix. By using a convolutional neural network (CNN) and a long short-term memory network (LSTM) to extract local and global features of the time series, the distance metric method is dynamically adjusted. The core idea is to optimize the similarity metric through a learnable parameter model, thereby improving the ability to capture complex time series features. The formula can be further extended to:

[0080] D(i,j) = ||f(A i ) - h(B j )||

[0081] where f(·) and h(·) are feature transformation functions learned through a deep neural network. They can map the original low-dimensional features to a high-dimensional feature space, enabling better measurement of the similarity between sequences in the new feature space. D(i,j) represents the distance metric between the two sequences at position (i,j), and ||·|| represents the Euclidean distance function.

[0082] (2) Optimize path selection

[0083] In the DTW path search, this method uses the backpropagation mechanism of deep learning to optimize the cumulative distance of dynamic programming. The recurrence formula for the cumulative distance is:

[0084] C(i,j) = D(i,j) + min(C(i - 1,j), C(i,j - 1), C(i - 1,j - 1))

[0085] Among them, C(i,j) represents the minimum cumulative distance from the point (0,0) to the position (i,j). C(i - 1,j) represents the cumulative cost in the vertical direction, that is, the path from the upper grid, corresponding to the case of deleting an element from sequence A; C(i,j - 1) represents the cumulative cost in the horizontal direction, that is, the path from the left grid, corresponding to the case of deleting an element from sequence B; C(i - 1,j - 1) represents the cumulative cost in the diagonal direction, that is, the path from the upper - left grid, corresponding to the case where two sequence elements match.

[0086] Through deep - learning optimization, the model can adaptively adjust the path - selection strategy, accurately capture the complex non - linear changes in the time series. To enhance the robustness of the algorithm, a path - optimization method based on regularization is designed. By introducing a smoothing regularization term R(C), its objective function is:

[0087] Loss = C(i,j)+λR(C), R(C)=λ(∑C(i + 1,j)-C(i,j) 2 +∑C(i,j + 1)-C(i,j) 2 )

[0088] Among them, λ is the regularization parameter, and R(C) is used to limit the volatility of path selection and improve the generalization ability of the model.

[0089] Backward - gradient update method. This method first determines an optimal alignment path, and then calculates the gradient only on this fixed path, that is, calculates the backward gradient at the points on the fixed alignment path. The formula is as follows:

[0090]

[0091] Among them, x represents the source time series, y represents the time series to be aligned, represents the partial derivative of x, and dtw 2 represents the square of the dtw function, which can ensure that the gradient is positive and smooth, avoid the high - variance problem of the traditional SoftDTW method, and significantly improve the calculation efficiency.

[0092] (3) FastDTW (FastDTW) optimization

[0093] To solve the problem of the high computational complexity of the traditional DTW algorithm, this method adopts a multi - scale approximation method, quickly calculates the initial alignment path at a low resolution, and gradually optimizes the path at a high resolution. Finally, the computational complexity is reduced from O(n 2)Reduce it to O(n log n). At the same time, utilize the parallel computing and GPU acceleration mechanisms of the deep learning framework to further improve the computational efficiency of path search, achieve efficient processing of large-scale data, and dynamically adjust the path optimization strategy at different resolutions by weighting the feature importance through the combination of deep learning modules.

[0094] Specifically, low resolution and high resolution refer to the sampling rate or density of time series data. Quickly find a rough alignment at low resolution. Low resolution: The sequence length is shorter (usually < 100 time points), and the data points are sparse. High resolution: The sequence length is longer (usually > 100 time points), and the data points are dense.

[0095] Low-resolution initial alignment: Downsample the original sequence (e.g., take 1 point every 4 points); calculate the DTW path on the downsampled sequence; quickly obtain a rough alignment result; significantly reduce the computational complexity.

[0096] Medium-resolution optimization: Increase the sampling rate (e.g., take 1 point every 2 points); perform local search based on the low-resolution result; optimize the alignment path within a small window; balance computational efficiency and alignment accuracy.

[0097] High-resolution refinement: Perform the final optimization at the original resolution; only perform local adjustments near the determined path; achieve precise sequence alignment; ensure the accuracy of the final result.

[0098] S3, Input the preprocessed data and the collected historical data into the deep learning-optimized DTW model to obtain the similarity of time series.

[0099] Collect historical data. In bacteria detection, the historical data is the data of different bacteria in different states; in climate prediction, the historical data is historical climate data; in financial market analysis, the historical data is historical market volatility patterns.

[0100] S4, Based on the similarity of time series, use the K-nearest neighbor algorithm to classify the preprocessed data to complete the time series similarity evaluation.

[0101] Combine the deep learning-enhanced dynamic time warping (DTW) algorithm with the K-nearest neighbor (KNN) classification algorithm for time series classification, which specifically includes the following:

[0102] S41, Select k most neighboring samples: Use the deep learning-enhanced DTW to calculate the similarity of time series, and select the k historical samples most similar to the input time series as the most neighboring samples.

[0103] S42, Weight Assignment: Depending on the similarity output by the deep learning model, the k nearest neighbor samples are weighted, and the weight calculation formula is as follows:

[0104]

[0105] where w α is the weight of the α-th nearest neighbor sample, and dist(A, B α ) is the similarity distance between the current sequence A and the nearest neighbor sample B α , which is calculated by the DTW algorithm controlled by the deep learning model.

[0106] S42, Classification Decision: Classification decision is made using the weighted voting of the nearest neighbor samples, and the formula is as follows:

[0107]

[0108] where c represents the sequence category, K represents the set of k nearest neighbor samples, represents the predicted category, w m represents the weight of the m-th nearest neighbor sample, y m represents the category label of the m-th nearest neighbor sample, Π is an indicator function that returns 1 when the condition is met and 0 otherwise. arg max represents the independent variable when the function reaches the maximum value, and here it represents finding the category with the highest score.

[0109] The time series similarity evaluation method based on DTW and K-nearest neighbor algorithm is applicable to the analysis of complex time series data in multiple fields, including the following scenarios:

[0110] (1) Bacterial Detection: By analyzing the time series of the surface mechanical expression (or direct surface tension) of bacteria, the DTW algorithm can be used to calculate and analyze the property similarity of different bacteria in different states, and combined with KNN classification to accurately identify the bacterial type.

[0111] (2) Climate Prediction: Compare the similarity between a certain time-temperature sequence and historical climate data, and use the KNN algorithm to predict future climate types (such as sunny days, rainy days, etc.) to improve the accuracy and efficiency of meteorological prediction.

[0112] (3) Financial Market Analysis: Analyze the time series of stock price fluctuations, use the DTW algorithm to find similar historical market fluctuation patterns, and use KNN classification to predict future market trends to help investors formulate more scientific investment strategies.

[0113] Example 1: Bacterial Detection Based on Liquid Gating Technology and Time Series Classification Algorithm.

[0114] Data Acquisition:

[0115] In this embodiment, a liquid-gated composite membrane is used to regulate the pressure change during the high-pressure fluid transmission process, and the generated pressure-time series data is used to characterize the surface properties of bacteria. Some of the collected original time series data is as Figure 5 shown, which shows the time series fluctuation characteristics under experimental conditions.

[0116] Data processing:

[0117] To ensure the stability and comparability of the time series data, the original data needs to be standardized and normalized to eliminate data differences caused by different experimental conditions and equipment. The processing methods include:

[0118] Data standardization: Convert the time series data to zero mean and unit variance to ensure the consistency of the data distribution.

[0119] Noise processing: Remove sudden outliers through the moving average method, and use a deep learning preprocessing module to further reduce the noise interference in the data.

[0120] The unprocessed data is as Figure 6 shown, which contains obvious noise and data fluctuations; while the preprocessed data is as Figure 7 shown, and its degree of dispersion is significantly reduced, and the characteristic trend is clearer.

[0121] Application of the Dynamic Time Warping (DTW) algorithm:

[0122] After the preprocessing is completed, the dynamic time warping (DTW) algorithm optimized by combining deep learning is used for the similarity evaluation of time series data. Combining with the deep learning module, this algorithm extracts local and global features of the time series layer by layer and dynamically adjusts the path search strategy to ensure efficient alignment in complex data scenarios.

[0123] In this embodiment, this algorithm is used to process time series data with high noise to accurately calculate the similarity between the target sequence and historical data. The specific operation process includes: loading the standardized time series data, inputting it into the DTW algorithm module optimized by combining deep learning, calculating the similarity matrix through path alignment, and passing the result to the classification module for subsequent classification and analysis. The whole process ensures the accuracy of similarity evaluation and provides stable and efficient input for the classification stage at the same time.

[0124] Classification by the K-Nearest Neighbor (KNN) algorithm:

[0125] After the DTW algorithm calculation is completed, the KNN algorithm is further used to classify and analyze bacterial samples. The specific steps include:

[0126] Proximity sample selection: Based on DTW similarity calculation, select the k historical samples closest to the current detected sample as references.

[0127] Weight assignment: Combine the feature importance extracted by the DTW deep learning module to assign weights to each proximity sample, ensuring that samples with high similarity have a greater impact on classification.

[0128] Classification voting: Determine the classification result of the current sample through majority voting, and finally output the bacterial species and their corresponding characteristic information.

[0129] Detection results and advantages:

[0130] According to the classification results of the DTW-KNN algorithm optimized by deep learning, output the surface characteristics of the bacteria and their corresponding time series pressure change curves. The specific advantages include:

[0131] High precision: Combine the DTW algorithm of deep learning and the optimized KNN classifier. Compared with the ordinary DTW algorithm, the classification recognition accuracy is significantly improved. As Figure 8 shown, the classification accuracy of the DTW algorithm optimized by deep learning reaches 79% under noise-free conditions, while the traditional DTW-KNN algorithm is only 39%. Under noisy conditions, the DTW algorithm optimized by deep learning still maintains an accuracy of 42%, far higher than the 25% of the traditional DTW-KNN algorithm. Generally, the DTW algorithm optimized by deep learning shows stronger robustness in a noisy environment.

[0132] Automation: It can automatically complete data processing, sequence alignment, and classification analysis, minimizing the possibility of human intervention. Through the optimization of the FastDTW algorithm, as Figure 9 shown, in terms of running time, the FastDTW-KNN algorithm only needs 30 seconds, while the traditional DTW-KNN algorithm needs 65 seconds, and the LSTM algorithm is as high as 247 seconds. It shows that this embodiment significantly reduces the running time while ensuring high precision, improving the computing efficiency.

[0133] Scalability: As the historical database continues to expand, it can support more diverse bacterial sample detection requirements and adapt to complex non-linear time series analysis scenarios.

[0134] Table 1: Comparative evaluation of the effects of using different algorithms to process bacterial detection data.

[0135]

[0136] Example 2: Climate prediction and planting decision-making in agriculture.

[0137] This embodiment targets the historical meteorological data of a certain region from 2000 to 2020 in summer. The data includes multi-dimensional meteorological features such as temperature, rainfall, and wind speed. After dimensionality reduction and sequence matching using the fast DTW (Dynamic Time Warping optimized by deep learning) algorithm, the meteorological conditions on a specific future date are predicted.

[0138] Data preprocessing:

[0139] Data collection: Collect the summer meteorological data of a certain place in the past 20 years, including climate data with multi-dimensional features such as continuous temperature, wind speed, rainfall, and humidity for 24 hours a day, as Figure 10 shown.

[0140] Dimensionality reduction processing: Use LDA (Linear Discriminant Analysis) to reduce the dimensionality of high-dimensional data and extract key features. For example, reduce the original multi-dimensional climate data to 2 - 3 main features to simplify subsequent calculations.

[0141] Data standardization: Normalize and standardize the data to eliminate the influence of different eigenvalue ranges and improve the alignment accuracy.

[0142] Fast DTW path alignment:

[0143] Based on the historical meteorological time series, the fast DTW algorithm performs dynamic similarity matching on historical data through multi-scale path search, and calculates the most similar time series path between the target date and the historical data.

[0144] Introduce a feature weighting mechanism optimized by deep learning to dynamically adjust the alignment weights, further improving the accuracy and robustness of path matching.

[0145] Climate prediction output:

[0146] According to the most similar feature path in the historical time series, output key information such as temperature, rainfall, and wind speed on a specific future date.

[0147] Automatically generate a visual prediction curve of the time series for analyzing and interpreting the prediction results.

[0148] Detection results. As Figure 11 shown, the detection results of this embodiment on the experimental dataset are as follows:

[0149] Prediction accuracy: The fast DTW algorithm achieved a prediction accuracy of 79% in the dataset, significantly higher than 39% of the traditional DTW and 33% of the LSTM. This indicates that after being optimized by deep learning, the fast DTW can more accurately match complex time series.

[0150] Running time: The running time of Fast DTW is 30 hours, which is only 46% of traditional DTW (65 hours) and far lower than 247 hours of LSTM. Fast DTW has significantly improved time efficiency in large-scale data processing and is suitable for practical application scenarios.

[0151] Advantages of the method in this embodiment:

[0152] High precision: Fast DTW combines a path optimization strategy of deep learning, effectively capturing local non-linear changes in time series and improving the accuracy of climate prediction.

[0153] High efficiency: The algorithm reduces the time complexity of path alignment from O(n 2 ) of traditional DTW to O(n log n), greatly reducing the computational overhead.

[0154] Robustness: In the dimensionality-reduced time series data with high noise, Fast DTW still maintains stable prediction performance, demonstrating its adaptability to non-linear and high-noise data.

[0155] Scalability: Fast DTW is applicable to larger-scale historical meteorological databases and can be effectively extended to more diverse climate prediction tasks, such as multi-season or cross-year prediction.

[0156] Table 2: Comparative evaluation of the effects of processing meteorological data using different algorithms.

[0157]

[0158] The method in this embodiment has a DTW core enhanced by deep learning: By deep learning, the alignment path of DTW is made dynamic and learnable, thus adapting to different types of time series data and solving the adaptability problem of traditional DTW in non-linear complex data.

[0159] The method in this embodiment fuses KNN classification with deep learning features: Combining deep learning feature extraction and KNN classification algorithm, and using a similarity weighting mechanism to achieve efficient classification of complex time series.

[0160] The method in this embodiment integrates performance optimization with Fast DTW: Combining the Fast DTW algorithm and multi-scale feature extraction of deep learning models, improving the computational efficiency while ensuring the application performance of the model on large-scale data sets.

[0161] The method in this embodiment is easy to interpret and apply: The classification result is not only a single output, but also provides a similarity comparison with historical data, having good interpretability, facilitating users to refer to and verify in detection, and even adjusting the result in combination with artificial experience.

[0162] The method of this embodiment has strong scalability: with the accumulation of historical data, the classification and prediction accuracy of the present invention will continue to improve, and higher accuracy and reliability can be achieved by continuously optimizing the database.

[0163] The method of this embodiment not only overcomes the limitations of traditional methods in non-linear data processing and long-time series analysis, but also provides transparent and interpretable classification results, and is applicable to bacterial detection data, climate prediction data, financial data, and bioinformatics data analysis.

[0164] See Figure 12 As shown, the present invention also discloses a time series similarity evaluation device based on the DTW and K-nearest neighbor algorithms, including:

[0165] A data acquisition and preprocessing module 1201, configured to collect target time series data in real time and preprocess the data;

[0166] A model construction module 1202, configured to construct a deep learning model including a DTW Function module, a DTW Layer, and an optimizer; the DTW Function module is a differentiable neural network component based on the DTW algorithm, and realizes forward path calculation and reverse gradient calculation; the DTW Layer encapsulates the DTW algorithm into a neural network layer and extracts time series features layer by layer; the optimizer updates the parameters of the deep learning model;

[0167] A time series similarity acquisition module 1203, configured to input the preprocessed data and the collected historical data into the deep learning model to obtain the similarity of the time series;

[0168] A K-nearest neighbor algorithm category prediction module 1204, configured to classify the preprocessed data using the K-nearest neighbor algorithm based on the similarity of the time series to complete the time series similarity evaluation.

[0169] The specific implementation of the time series similarity evaluation device based on the DTW and K-nearest neighbor algorithms is the same as that of the time series similarity evaluation method based on the DTW and K-nearest neighbor algorithms, and will not be repeated in this embodiment.

[0170] The above is only the specific implementation manner of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection of the present invention.

Claims

1. A time series similarity evaluation method based on DTW and K nearest neighbor algorithm, characterized in that: The steps include: S1, collects target time series data in real time and preprocesses the data; S2, constructing a deep learning model including a DTW Function module, a DTW Layer and an optimizer; the DTW Function module is a differentiable neural network component based on the DTW algorithm, which implements forward path calculation and reverse gradient calculation; the DTW Layer encapsulates the DTW algorithm into a neural network layer, and extracts time series features layer by layer; the optimizer updates the parameters of the deep learning model; S3, input the preprocessed data and the collected historical data into the deep learning model to obtain the similarity of the time series; S4, based on the similarity of time series, uses the K nearest neighbor algorithm to classify the preprocessed data and complete the time series similarity evaluation.

2. The time series similarity evaluation method based on DTW and K nearest neighbor algorithm according to claim 1 is characterized in that: The preprocessing includes data standardization and normalization, noise processing and missing value filling; the noise processing adopts the sliding average method; the missing value filling adopts linear interpolation and automatic filling technology.

3. The time series similarity evaluation method based on DTW and K nearest neighbor algorithm according to claim 1 is characterized in that: The optimizer is the AdamW optimizer.

4. The time series similarity evaluation method based on DTW and K nearest neighbor algorithm according to claim 3 is characterized in that: The reverse gradient calculation is specifically as follows: Determine an optimal alignment path as a fixed path, and then calculate the reverse gradient only on this fixed path; the reverse gradient calculation formula is expressed as: Among them, x represents the source time series; y represents the time series to be aligned; Indicates partial derivative with respect to x; dtw 2 Indicates that the dtw function is squared.

5. The time series similarity evaluation method based on DTW and K nearest neighbor algorithm according to claim 1 is characterized in that: The objective function of the deep learning model is: Loss=C(i,j)+λR(C) Among them, Loss represents the objective function; C(i,j) represents the minimum cumulative distance from point (0,0) to position (i,j); λ represents the regularization parameter; R(C) represents the smoothing regularization term.

6. The time series similarity evaluation method based on DTW and K nearest neighbor algorithm according to claim 5 is characterized in that: The recursive formula of the minimum cumulative distance is expressed as: C(i,j)=D(i,j)+min(C(i-1,j),C(i,j-1),C(i-1,j-1)) Where D(i,j) is a trainable distance matrix, which represents the distance measure between sequence A and sequence B at position (i,j); min(·) means taking the minimum value; C(i-1,j) represents the cumulative cost in the vertical direction; C(i,j-1) represents the cumulative cost in the horizontal direction; C(i-1,j-1) represents the cumulative cost in the diagonal direction.

7. The time series similarity evaluation method based on DTW and K-nearest neighbor algorithm according to claim 6 is characterized in that: The trainable distance matrix is ​​expressed as: D(i,j)=||f(A i )-h(B j )|| Where f(·) and h(·) represent the feature transformation functions learned through deep neural networks; A i represents the i-th point in sequence A; B j represents the j-th point in sequence B; ||·|| represents the Euclidean distance function.

8. The time series similarity evaluation method based on DTW and K nearest neighbor algorithm according to claim 1 is characterized in that: The deep learning model uses a multi-scale approximation method for FastDTW optimization.

9. The time series similarity evaluation method based on DTW and K-nearest neighbor algorithm according to claim 1, characterized in that: The method of classifying the preprocessed data using the K nearest neighbor algorithm based on the similarity of the time series includes the following steps: S41, based on the similarity of the time series, select k historical data that are most similar to the time series of the preprocessed data as the nearest neighbor samples; S42, weights are assigned to the k nearest neighbor samples, and the weight calculation formula is as follows: Among them, w α represents the weight of the αth nearest neighbor sample; dist(A,B α ) represents the current sequence A and the nearest sample B α The time series similarity between S42, weighted voting is performed on the k nearest neighbor samples to obtain the predicted category of the preprocessed data; the formula is as follows: Where c represents the type of sequence; arg max represents the independent variable when the function reaches the maximum value; K represents the set of k nearest neighbor samples; Indicates the predicted category; w m represents the weight of the mth nearest neighbor sample; y m represents the category label of the mth nearest neighbor sample; Π represents the indicator function, which returns 1 when the condition is met, otherwise it returns 0.

10. A time series similarity evaluation device based on DTW and K nearest neighbor algorithm, comprising the following: Data collection and preprocessing module, used to collect target time series data in real time and preprocess the data; A model building module is used to build a deep learning model including a DTW Function module, a DTW Layer and an AdamW optimizer; the DTW Function module is a differentiable neural network component based on the DTW algorithm, which implements forward path calculation and reverse gradient calculation; the DTW Layer encapsulates the DTW algorithm into a neural network layer and extracts time series features layer by layer; the optimizer updates the deep learning model parameters; The time series similarity acquisition module is used to input the preprocessed data and the collected historical data into the deep learning model to obtain the similarity of the time series; The K-nearest neighbor algorithm category prediction module is used to classify the preprocessed data based on the similarity of time series and complete the time series similarity evaluation using the K-nearest neighbor algorithm.

Citation Information

Cited By

  • Sequence distance measurement method and device based on time sequence synchronous prediction

    CN121115015A

  • Pilot ability assessment method based on dynamic time warping and hierarchical clustering

    CN121146625A