Neural architecture search technology and lightweight Transform-based time sequence classification method
By combining neural architecture search technology and lightweight Transformer model, a model structure with decoupling of spatiotemporal features is constructed, and the CMA algorithm and adaptive global pruning technology are optimized, the problem of complex time series data processing in the existing technology is solved, and efficient and lightweight time series classification is achieved.
Patent Information
- Application Number
- CN202510154688.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
When processing complex time series data, it is difficult for the prior art to automatically adapt to different data characteristics, and the model is difficult to achieve lightweight while maintaining high classification performance, high computational complexity, and difficult to apply in resource-constrained environments.
The time series classification method based on neural architecture search technology and lightweight Transformer is adopted to build a model structure through the principle of spatiotemporal feature decoupling, and model optimization is carried out in combination with CMA algorithm and adaptive global pruning technology to realize automatic search and hyperparameter optimization of the model.
It realizes that the model maintains high classification performance while significantly reducing the computational complexity, can automatically adapt to the data characteristics of different time series, and is suitable for resource-constrained environments, improving the adaptability and reliability of the model.
Smart Images

Figure CN120067806A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, in particular to a time series classification method based on neural architecture search technology and lightweight Transformer. Background Art
[0002] Time series classification is an important issue in the fields of data analysis and machine learning, and is widely applied in multiple fields such as financial forecasting, meteorological analysis, industrial production monitoring, etc. With the advent of the big data era, the scale and complexity of time series data have been continuously increasing, and traditional classification methods face many challenges when dealing with these complex data.
[0003] In recent years, deep learning technology, especially the Transformer model, has achieved remarkable success in dealing with sequence data. However, there are still some problems in directly applying the Transformer to time series classification tasks. First of all, the computational complexity of the standard Transformer model is relatively high, and it is difficult to process long sequence data. Secondly, the structural design of the Transformer model is mainly aimed at natural language processing tasks and does not fully consider the special properties of time series data. In addition, existing Transformer models often require a large amount of labeled data and computational resources to achieve ideal performance, which is difficult to meet in many actual application scenarios.
[0004] On the other hand, neural architecture search (NAS) technology has shown great potential in automatically designing neural network structures. However, most existing NAS methods focus on image processing or natural language processing tasks and lack consideration for time series classification problems. In addition, traditional NAS methods usually require a large amount of computational resources and time and are difficult to apply in resource-constrained environments.
[0005] The closest prior art attempts to combine lightweight Transformer with NAS technology to solve the time series classification problem. However, these methods still have some obvious deficiencies. First of all, they do not fully consider the spatio-temporal characteristics of time series data, resulting in poor performance of the model in capturing long-term dependencies and local patterns. Secondly, existing methods lack flexibility and adaptability in the model search and optimization process and are difficult to adapt to different types of time series data. Moreover, the balance between model lightweighting and performance of these methods is not ideal enough, and often significantly reduces the classification accuracy while reducing the model complexity.
[0006] In addition, when dealing with complex time series data such as high-dimensional, noisy, and data-imbalanced data, existing technologies do not perform stably enough. They usually require a large amount of manual intervention to adjust the model structure and hyperparameters, which not only increases the usage cost but also limits the universality of the method. At the same time, existing methods also face challenges in processing real-time data streams and performing online learning, and it is difficult to meet some application scenarios with high real-time requirements.
[0007] In view of the above problems, there is an urgent need for a time series classification method that can automatically adapt to the characteristics of different time series data, achieve model lightweight while maintaining high classification performance, and can efficiently process complex data. Summary of the Invention
[0008] The present invention provides a time series classification method based on neural architecture search technology and lightweight Transformer, aiming to solve the problems existing in the prior art. Specifically, the technical problems to be solved by the present invention include: how to design a model structure that can automatically adapt to the characteristics of different time series data; how to achieve model lightweight while maintaining high classification performance; how to efficiently process complex time series data with long sequences, high dimensions, and large noise; and how to achieve efficient model search and optimization in resource-constrained environments.
[0009] The present invention proposes a time series classification method based on neural architecture search technology and lightweight Transformer, including:
[0010] An acquisition step, including:
[0011] Obtain a time series data set and predefined classification labels;
[0012] A preprocessing step, including:
[0013] Based on the time series data set, perform data cleaning, missing value filling, and normalization processing;
[0014] According to the processed time series data, generate a two-dimensional data set including time features, weather features, and temperature features;
[0015] A model construction step, including:
[0016] Based on the spatio-temporal feature decoupling principle, construct a lightweight Transformer model structure;
[0017] According to the lightweight Transformer model structure, generate a neural network model including a time feature extraction module, a spatial feature extraction module, and a spatial residual extraction module;
[0018] A model optimization step, including:
[0019] Based on the CMA algorithm and adaptive global pruning technology, perform structure search and hyperparameter optimization on the neural network model;
[0020] According to the optimization results, obtain the optimal lightweight Transformer model structure;
[0021] The training steps include:
[0022] Based on the two-dimensional dataset and the optimal lightweight Transformer model structure, perform model training;
[0023] According to the training results, obtain the trained model for time series classification;
[0024] The classification steps include:
[0025] Based on the trained model, perform feature extraction and classification prediction on the input time series data;
[0026] According to the prediction results, output the classification label of the time series.
[0027] Preferably, the preprocessing step specifically includes:
[0028] Interpolate and fill the missing values in the time series dataset;
[0029] Use the min-max normalization method to normalize the data;
[0030] Generate a three-dimensional feature vector for each time step, including the current moment, the corresponding weather data, and the temperature at this time node.
[0031] Preferably, the lightweight Transformer model structure includes:
[0032] Multiple spatio-temporal feature decoupling modules, each module includes a time feature extraction sub-module, a spatial feature extraction sub-module, and a spatial residual extraction sub-module;
[0033] The time feature extraction sub-module maps the features of the input sequence to time feature encoding through time fusion;
[0034] The spatial feature extraction sub-module and the spatial residual extraction sub-module are respectively implemented through standard Transformer layers and fully connected layers, and are combined through a residual connection method.
[0035] Preferably, the model optimization step specifically includes:
[0036] Set the search space, including the number of modules, layers, width, and activation function of the Transformer encoder;
[0037] Generate candidate model structures using the CMA algorithm;
[0038] Apply the adaptive global pruning technique to each candidate model structure and dynamically adjust the model parameters;
[0039] Evaluate the pruned model structures based on the performance on the validation set;
[0040] Iteratively update the search strategy until the optimal model structure is found.
[0041] Preferably, the adaptive global pruning technique includes:
[0042] Set importance parameters for the attention weights in the model;
[0043] Update the importance parameters through the gradient backpropagation algorithm;
[0044] Penalize the gradients based on the parameter importance to achieve automatic pruning of unimportant weights;
[0045] Dynamically adjust the pruning layer number according to the validation accuracy.
[0046] Preferably, the training step adopts a two-stage training strategy:
[0047] The first stage: Train the initial Transformer structure;
[0048] The second stage: Select the best structure using cross-validation and perform fine-tuning.
[0049] Preferably, it further includes a computational complexity optimization step:
[0050] Based on the lightweight Transformer model structure, control the computational complexity within O(n*d), where n is the time series length and d is the input feature dimension.
[0051] Preferably, the model optimization step further includes:
[0052] Further optimize the Transformer model by combining the local search strategy and the gradient descent method;
[0053] Use the L1 regularization method to prune unimportant model components and balance the pruning and search tasks.
[0054] Preferably, it further includes an automatic learning strategy step:
[0055] Automatically adjust the model search strategy based on the performance results of different training data;
[0056] Dynamically update the search space and optimization objective to adapt to the characteristics of different time series data.
[0057] Preferably, the classification step further includes:
[0058] Performing real-time feature rearrangement on the input time series data;
[0059] Using the trained lightweight Transformer model to extract spatio-temporal features;
[0060] Generating a classification probability distribution through a fully connected layer and a softmax function;
[0061] Selecting the class with the highest probability as the final classification result.
[0062] By innovatively combining neural architecture search technology and lightweight Transformer models, the present invention proposes a comprehensive and efficient time series classification solution. This method can not only automatically design a model structure suitable for specific tasks, but also significantly reduce the computational complexity while maintaining high classification performance.
[0063] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0064] First, the lightweight Transformer model structure with decoupled spatio-temporal features proposed by the present invention can effectively capture long-term dependencies and local patterns in time series data. This structural design not only improves the model's ability to understand complex time series, but also greatly reduces the computational complexity, enabling the model to efficiently process long sequence data. For example, in financial time series prediction tasks, this method can simultaneously consider long-term economic trends and short-term market fluctuations, thus making more accurate predictions.
[0065] Second, the model optimization strategy of the present invention combining the CMA algorithm and adaptive global pruning technology realizes the joint optimization of the model structure and hyperparameters. This strategy not only improves the search efficiency, but also enables effective lightweighting while maintaining the model performance. In practical applications, this means that high-performance time series classification models can be deployed in resource-constrained environments (such as mobile devices or embedded systems).
[0066] Furthermore, the automatic learning strategy of the present invention can dynamically adjust the model search strategy according to the performance results of different training data. This adaptive ability enables this method to handle various types of time series data, from simple univariate time series to complex multivariate and high-dimensional time series. For example, in meteorological prediction tasks, this method can automatically adapt to the meteorological data characteristics of different regions and different time scales, providing more accurate prediction results.
[0067] In addition, the innovative preprocessing and feature rearrangement techniques of the present invention improve the utilization efficiency of the model for the original time series data. This not only enhances the robustness of the model but also enables the model to better handle complex situations such as high noise and data imbalance. In applications such as industrial production monitoring, this ability can help the model more accurately identify abnormal patterns and improve the accuracy of fault prediction.
[0068] Finally, by comprehensively applying various optimization techniques such as L1 regularization and gradient clipping, the present invention achieves a multi-faceted balance among model performance, computational efficiency, and generalization ability. This enables the method of the present invention to not only perform excellently in specific tasks but also have good transfer learning ability and can quickly adapt to new time series classification tasks.
[0069] Generally speaking, the method provided by the present invention effectively solves the problems existing in the prior art through innovative model design, optimization strategies, and adaptive mechanisms. It not only improves the accuracy and efficiency of time series classification but also greatly enhances the adaptability and reliability of the method. This method provides a new technical path for the field of time series classification and is expected to bring significant application value in multiple fields such as financial forecasting, meteorological analysis, and industrial monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is the overall flowchart of the present invention.
[0071] Figure 2 is the data preprocessing diagram of the present invention.
[0072] Figure 3 is the logical block diagram of the model construction steps of the present invention.
[0073] Figure 4 is the logical block diagram of the model optimization steps of the present invention.
[0074] Figure 5 is the logical block diagram of the model training and classification steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0075] Please refer to the attached Figures 1-5 , the present invention provides a time series classification method based on neural architecture search technology and lightweight Transformer. By innovatively combining neural architecture search and lightweight Transformer technology, this method effectively solves the challenges faced by traditional time series classification methods in dealing with complex data. The following will describe the detailed implementation manners of the present invention.
[0076] As described in claim 1, the method of the present invention includes an acquisition step, a preprocessing step, a model construction step, a model optimization step, a training step, and a classification step. These steps form a complete time series classification process, and each step is carefully designed to achieve the optimal classification effect.
[0077] First, in the acquisition step, the method acquires a time series data set and predefined classification labels. This data may come from various fields, such as finance, meteorology, or industrial production, etc. For example, in the financial field, the time series data may be the historical data of stock prices, and the classification labels may be "rise", "fall", or "flat".
[0078] Next, the preprocessing step cleans, fills, and normalizes the original data. This step is crucial because high-quality input data is the basis for model performance. The present invention adopts advanced data preprocessing techniques, including but not limited to filling missing values by interpolation methods and the min-max normalization method. For example, for missing values, methods such as linear interpolation or forward filling can be used. The normalization process can map the data to the [0,1] interval, which helps to improve the training efficiency and generalization ability of the model.
[0079] In the model construction step, the present invention constructs an innovative lightweight Transformer model structure based on the principle of spatio-temporal feature decoupling. This structure includes a time feature extraction module, a spatial feature extraction module, and a spatial residual extraction module, which can effectively capture the time dependence and spatial correlation in time series data. For example, the time feature extraction module can learn long-term dependencies through the self-attention mechanism, while the spatial feature extraction module can capture the interactions between different features.
[0080] The model optimization step is a major innovation point of the present invention. It combines the CMA (Covariance Matrix Adaptation) algorithm and the adaptive global pruning technique to achieve the automatic search of the model structure and the optimization of hyperparameters. The CMA algorithm is an efficient global optimization method, especially suitable for dealing with complex non-linear optimization problems. For example, it can be used to search for optimal structural parameters such as the number of model layers and the number of attention heads. The adaptive global pruning technique realizes the lightweight of the model by dynamically adjusting the model parameters while maintaining high performance.
[0081] Preferably, in one embodiment of the present invention, the specific implementation of the CMA algorithm can be expressed as follows:
[0082]
[0083] Where m is the distribution mean, σ is the step size, w i is the weight, x i: λ is the sorted individual, p σ is the evolutionary path, c σ and d σ are the learning rate parameters.
[0084] In the training step, the method adopts a two-stage training strategy, making full use of cross-validation techniques to select the best model structure. This strategy not only improves the generalization ability of the model but also effectively prevents the occurrence of overfitting problems. For example, in the first stage, 80% of the data can be used for initial training, and the remaining 20% for validation. In the second stage, 5-fold cross-validation can be used to fine-tune the model parameters.
[0085] Finally, in the classification step, the trained model is used to classify and predict new time series data. The method of the present invention achieves efficient and accurate time series classification through optimized feature extraction and classification processes.
[0086] The preprocessing step of the present invention further refines the data processing flow. For missing value filling, the method preferably uses interpolation. For example, for linear interpolation, the following formula can be used:
[0087]
[0088] where x is the point to be interpolated, (x 1 , y 1 ) and (x 2 , y 2 ) are the known adjacent points. For data normalization, the method adopts the min-max normalization method, and its formula is:
[0089]
[0090] where x is the original value, x m in and x m ax are the minimum and maximum values of the data respectively.
[0091] This normalization method maps the data to the [0, 1] interval, which is beneficial to improving the training efficiency and stability of the model.
[0092] In addition, the present invention innovatively generates a three-dimensional feature vector for each time step, including the current moment, the corresponding weather data, and the temperature at that time node. This multi-dimensional feature representation greatly enhances the model's ability to understand time series data. For example, in meteorological prediction tasks, this feature representation can help the model better capture the relationship between temperature changes and weather conditions.
[0093] The lightweight Transformer model structure of the present invention is an innovative design. This structure contains multiple spatio-temporal feature decoupling modules, and each module is composed of a time feature extraction sub-module, a spatial feature extraction sub-module, and a spatial residual extraction sub-module. This modular design not only improves the interpretability of the model but also enhances its adaptability to different types of time series data.
[0094] The time feature extraction sub-module maps the features of the input sequence to time feature encodings through time fusion technology. This process can be represented by the following formula:
[0095] T = TimeEmbedding(X)·W T ,
[0096] where X is the input sequence, TimeEmbedding is the time embedding function, and W T is a learnable weight matrix. The spatial feature extraction sub-module and the spatial residual extraction sub-module are implemented through a standard Transformer layer and a fully connected layer respectively. The self-attention mechanism of the standard Transformer layer can be expressed as:
[0097]
[0098] where Q, K, and V are the query, key, and value matrices respectively, and d k is the dimension of the key. The spatial residual extraction sub-module is combined with the spatial feature extraction sub-module through a residual connection, which can be expressed as:
[0099] Y = LayerNorm(X + FFN(Attention(X))),
[0100] where FFN is the feed-forward neural network and LayerNorm is the layer normalization operation.
[0101] This innovative model structure design enables the present invention to effectively process complex time series data, achieving model lightweight while maintaining high performance. For example, in financial time series prediction tasks, the model can simultaneously capture the long-term trend of stock prices (through time feature extraction) and the correlation between different economic indicators (through spatial feature extraction), thereby making more accurate predictions.
[0102] From the above detailed description, it can be seen that the method of the present invention has significant advantages in time series classification tasks. It can not only effectively process complex time series data but also achieve efficient classification performance through an innovative model structure and optimization strategy. This method has broad application prospects in multiple fields such as financial forecasting, meteorological analysis, and industrial production monitoring.
[0103] The method of the present invention adopts an innovative strategy in the model optimization step, as described in claim 4. The core of this step lies in the combination of the CMA algorithm and the adaptive global pruning technique, achieving the automatic search of the model structure and the optimization of hyperparameters.
[0104] First, the method sets a flexible search space, including the number of modules, the number of layers, the width, and the activation function of the Transformer encoder. This comprehensive search space design allows the model to be optimized in multiple dimensions, so as to find the most suitable structure for a specific time series classification task. For example, in one embodiment, the number of modules can vary from 1 to 8, the number of layers can be searched from 1 to 6, and the width can be adjusted from 64 to 512. The selection of activation functions can include common functions such as ReLU, GELU, and Swish.
[0105] Next, the present invention uses the CMA algorithm to generate candidate model structures. As an efficient global optimization method, the CMA algorithm is particularly suitable for dealing with such complex non-convex optimization problems. Preferably, in one embodiment of the present invention, the update rule of the CMA algorithm can be expressed as:
[0106]
[0107] where m t is the distribution mean at time t, η m is the learning rate, w i is the weight, x i : λ are the sorted individuals. For each candidate model structure generated by the CMA algorithm, the method applies the adaptive global pruning technique to dynamically adjust the model parameters. This step is another innovation point of the present invention, which can significantly reduce the complexity of the model while maintaining the model performance. Specifically, the pruning process can be expressed as:
[0108] W pruned = W ⊙ M,
[0109] where is the original weight matrix, is the binary mask matrix, represents elementwise multiplication. The elements of the mask matrix are dynamically determined according to the importance of the weights, and the importance can be evaluated by the magnitude of the gradient.
[0110] In the model evaluation stage, the method evaluates the pruned model structure based on the performance on the validation set. The performance metrics here can be accuracy, F1 score, or a custom metric for a specific task. For example, in a financial time series prediction task, the Sharpe ratio can be used as an evaluation metric.
[0111] Finally, the method of the present invention continuously optimizes the model structure by iteratively updating the search strategy. This process will continue until the optimal model structure is found or the preset number of iterations is reached. Preferably, in one embodiment, the number of iterations can be set between 100 and 500, depending on the computing resources and time constraints.
[0112] The adaptive global pruning technique of the present invention is a key component in the model optimization process. This technique realizes the lightweight of the model by dynamically adjusting the model parameters while maintaining high performance.
[0113] First, the method sets importance parameters for the attention weights in the model. These importance parameters can be initialized to 1, indicating that all weights are considered important initially. Preferably, in one embodiment, the importance parameters can be expressed as:
[0114] I w = sigmoid(α w ),
[0115] where α w is a learnable parameter and I w is the corresponding importance parameter. Next, the importance parameters are updated through the gradient backpropagation algorithm. This process can be expressed as:
[0116]
[0117] where η is the learning rate and L is the loss function.
[0118] The method of the present invention innovatively penalizes the gradient based on the parameter importance to achieve automatic pruning of unimportant weights. Specifically, the gradient update can be expressed as:
[0119]
[0120] This mechanism ensures that important weights can be fully updated while unimportant weights are gradually pruned.
[0121] Finally, the method dynamically adjusts the pruning layer according to the validation accuracy. This is an innovative design that can maximize the pruning effect while ensuring the model performance. For example, in one embodiment, a precision threshold τ (such as 0.95) can be set. When the validation accuracy is lower than τ, the pruning layer is reduced; when the accuracy is higher than τ, the pruning layer is increased.
[0122] The training step of the present invention adopts an innovative two-stage training strategy. This strategy makes full use of the cross-validation technique and effectively improves the generalization ability of the model.
[0123] In the first stage, the method trains the initial Transformer structure. The goal of this stage is to quickly obtain a baseline model. Preferably, in one embodiment, 80% of the data can be used for training, and the remaining 20% for validation. The training process can use the Adam optimizer, the learning rate can be set to 1e-4, and the batch size can be adjusted according to the data scale, usually between 32 and 256.
[0124] In the second stage, cross-validation is used to select the best structure and perform fine-tuning. The present invention preferably adopts 5-fold cross-validation, which can make full use of limited data and at the same time provide stable performance estimation. In the training of each fold, an early stopping strategy can be used to prevent overfitting. For example, the patience parameter can be set to 10, that is, if the performance of the validation set does not improve for 10 consecutive epochs, the training will stop.
[0125] The advantage of this two-stage strategy is that it can quickly obtain a usable model and improve the generalization ability of the model through fine cross-validation. Especially when dealing with complex time series data, this strategy can effectively balance training efficiency and model performance.
[0126] The present invention also includes an innovative computational complexity optimization step. The goal of this step is to control the computational complexity of the lightweight Transformer model at O(n·d), where n is the time series length and d is the input feature dimension. The computational complexity of traditional Transformer models is usually O(n 2 ·d), and the main bottleneck lies in the self-attention mechanism. The present invention significantly reduces the computational complexity through innovative model design. Specifically, the method adopts the following techniques:
[0127] 1. Sparse attention mechanism: Instead of calculating the attention between all time steps, only the time steps within a local window are concerned. This can reduce the complexity from O(n 2 ) to O(n·w), where w is a fixed window size.
[0128] 2. Linear attention: Use the kernel trick to linearize the attention calculation. For example, the following formula can be used:
[0129] Attention(Q,K,V)=φ(Q)(φ(K) T V),
[0130] where φ is a kernel function, and ReLU etc. can be selected.
[0131] 3. Parameter sharing: Share some parameters between different layers to reduce the total number of parameters. Through the combination of these techniques, the present invention has successfully controlled the computational complexity at the O(n·d) level, greatly improving the computational efficiency of the model. This enables the method to process longer time series data while maintaining low computational resource requirements.
[0132] Preferably, in an embodiment of the present invention, for a time series with a length of 1000 and a feature dimension of 64, using an 8-head attention mechanism and a 6-layer Transformer structure, the inference time of the model can be controlled within 10 milliseconds (on a standard GPU). This efficient computational performance makes the method particularly suitable for real-time processing and large-scale data analysis scenarios.
[0133] From the above detailed description, it can be seen that the method of the present invention has significant innovations in model optimization, training strategies, and computational efficiency. These innovations not only improve the accuracy of time series classification but also greatly enhance the practicality and scalability of the model. Whether in the fields of financial data analysis, industrial production monitoring, or meteorological forecasting, etc., the method has demonstrated great application potential.
[0134] The method of the present invention introduces more innovative techniques in the model optimization step, as described in claim 8. These techniques further improve the performance and efficiency of the model, making the method have significant advantages in dealing with complex time series classification tasks.
[0135] First, the method combines a local search strategy and the gradient descent method to further optimize the Transformer model. This combined strategy makes full use of the advantages of the two methods, ensuring both the globality of the search and the efficiency of optimization. Specifically, the local search strategy can quickly explore neighboring candidate structures in the model structure space, while the gradient descent method can finely adjust the model parameters.
[0136] Preferably, in an embodiment of the present invention, the local search strategy can adopt the simulated annealing algorithm. The temperature scheduling of this algorithm can be expressed as:
[0137] T(t) = T 0 ·α t ,
[0138] where T 0 is the initial temperature, usually set to 1.0, α is the cooling rate, which can be set between 0.9 and 0.99, and t is the current iteration number. The simulated annealing algorithm allows the algorithm to have a certain probability of accepting suboptimal solutions during the search process, thus avoiding being trapped in local optima. At the same time, the method uses the L1 regularization method to prune unimportant model components, achieving a balance between pruning and search tasks. The L1 regularization term can be expressed as:
[0139]
[0140] where λ is the regularization coefficient and w i is the model weight. By adjusting the value of λ, the sparsity of the model can be controlled. Preferably, in one embodiment, λ can be set between 1e-5 and 1e-3, and the specific value can be determined by cross-validation.
[0141] This method combining local search, gradient descent, and L1 regularization not only improves the performance of the model but also effectively reduces the complexity of the model. For example, when dealing with financial time series data, this method can significantly reduce the number of model parameters while maintaining a high prediction accuracy, making the model more suitable for deployment in resource-constrained environments.
[0142] The present invention also introduces an innovative automatic learning strategy step. The core idea of this step is to automatically adjust the model search strategy according to the performance results of different training data. This adaptive ability enables the method to better adapt to different types of time series data, improving the generality and robustness of the method.
[0143] Specifically, this method dynamically updates the search space and optimization objective to adapt to the characteristics of different time series data. This process can be represented as a meta-learning problem:
[0144]
[0145] where θ represents the hyperparameters of the model, T represents a specific task, D T represents the data of this task, and L represents the loss function.
[0146] Preferably, in one embodiment of the present invention, Bayesian optimization can be used to implement this automatic learning strategy. Bayesian optimization uses a Gaussian process to model the relationship between hyperparameters and performance, and then uses an acquisition function to determine the next hyperparameter combination to try. Commonly used acquisition functions include Expected Improvement (EI) and Upper Confidence Bound (UCB).
[0147] A significant advantage of this automatic learning strategy is that it can continuously improve the performance of the model over time. For example, when dealing with multiple related but slightly different time series classification tasks, this strategy can gradually learn a general and efficient search strategy, thus quickly finding a high-performance model structure for new tasks.
[0148] The classification step of the present invention includes a series of innovative processing procedures, further improving the accuracy and efficiency of classification.
[0149] First, the method performs real-time feature rearrangement on the input time series data. The purpose of this step is to convert the original time series data into a form more suitable for processing by the Transformer model. Preferably, in one embodiment, the feature rearrangement can adopt a sliding window technique to split the long sequence into multiple overlapping subsequences. The window size and stride can be adjusted according to the specific task. For example, for time series data of length 1000, a sliding window of size 100 and stride 50 can be used.
[0150] Next, the method uses the trained lightweight Transformer model to extract spatio-temporal features. This process can be expressed as:
[0151] H = Transformer(X),
[0152] where X is the input time series data and H is the extracted feature. Then, the method generates a classification probability distribution through a fully connected layer and the softmax function. This process can be expressed as:
[0153] P = softmax(W·H + b),
[0154] where W and b are the weights and biases of the fully connected layer respectively.
[0155] Finally, the method selects the class with the highest probability as the final classification result. This can be expressed as:
[0156]
[0157] These steps together ensure the efficiency and accuracy of the model in various application scenarios, especially in providing excellent performance even under limited resources.
[0158] Preferably, in one embodiment of the present invention, a confidence threshold can be introduced to improve the reliability of classification. For example, a threshold τ (such as 0.8) can be set, and the classification result is only output when the highest probability is greater than τ, otherwise it is marked as "uncertain". This mechanism is particularly useful in some application scenarios with extremely high accuracy requirements, such as medical diagnosis or financial risk assessment.
[0159] From the above detailed description, it can be seen that the method of the present invention has significant innovations in aspects such as model optimization, automatic learning strategy, and classification process. These innovations not only improve the accuracy and efficiency of time series classification, but also greatly enhance the adaptability and reliability of the method. Whether in processing complex financial data, variable meteorological data, or in the field of industrial production monitoring, etc., this method has demonstrated powerful performance and broad application prospects.
[0160] The method of the present invention provides a comprehensive and efficient solution for time series classification tasks by comprehensively applying neural architecture search technology, lightweight Transformer models, innovative optimization strategies, and adaptive learning mechanisms. This method can not only handle complex time series data that is difficult to deal with by traditional methods, but also achieve lightweight and efficient computing of the model while maintaining high performance. This makes the present method particularly suitable for application scenarios that require real-time processing and large-scale data analysis, providing new ideas and technical support for the development of the time series classification field.
[0161] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A time series classification method based on neural architecture search technology and lightweight Transformer, characterized by: include: The acquisition steps include: Get a time series dataset and predefined classification labels; Preprocessing steps include: Based on the time series data set, perform data cleaning, missing value filling and normalization processing; According to the processed time series data, a two-dimensional data set including time characteristics, weather characteristics and temperature characteristics is generated; Model building steps include: Based on the principle of decoupling of spatiotemporal features, a lightweight Transformer model structure is constructed; According to the lightweight Transformer model structure, a neural network model including a temporal feature extraction module, a spatial feature extraction module and a spatial residual extraction module is generated; Model optimization steps include: Based on the CMA algorithm and the adaptive global pruning technology, the structure search and hyperparameter optimization of the neural network model are performed; According to the optimization results, the optimal lightweight Transformer model structure is obtained; The training steps include: Based on the two-dimensional data set and the optimal lightweight Transformer model structure, perform model training; According to the training results, a training completion model for time series classification is obtained; The classification steps include: Based on the trained model, feature extraction and classification prediction are performed on the input time series data; According to the prediction results, the classification label of the time series is output.
2. The method according to claim 1, characterized in that The pre-processing step specifically includes: Performing interpolation to fill missing values in the time series data set; The data were normalized using the minimum-maximum normalization method; For each time step, a three-dimensional feature vector is generated, which contains the current time, the corresponding weather data, and the temperature at that time node.
3. The method according to claim 1, characterized in that The lightweight Transformer model structure includes: Multiple spatiotemporal feature decoupling modules, each module includes a temporal feature extraction submodule, a spatial feature extraction submodule and a spatial residual extraction submodule; The temporal feature extraction submodule maps the features of the input sequence to the temporal feature code through temporal fusion; The spatial feature extraction submodule and the spatial residual extraction submodule are implemented by a standard Transformer layer and a fully connected layer respectively, and are combined by a residual connection method.
4. The method according to claim 1, characterized in that: The model optimization step specifically includes: Set the search space, including the number of modules, layers, width, and activation function of the Transformer encoder; Generate candidate model structures using the CMA algorithm; Apply adaptive global pruning techniques to each candidate model structure to dynamically adjust model parameters; Evaluate the pruned model structure based on the validation set performance; The search strategy is updated iteratively until the optimal model structure is found.
5. The method according to claim 4, characterized in that The adaptive global pruning technology includes: Set importance parameters for attention weights in the model; Update the importance parameters through the gradient backpropagation algorithm; Penalize gradients based on parameter importance to automatically prune unimportant weights; The number of pruning layers is dynamically adjusted according to the verification accuracy.
6. The method according to claim 1, characterized in that The training step adopts a two-stage training strategy: Phase 1: Training the initial Transformer structure; Phase 2: Use cross-validation to select the best structure and perform fine-tuning.
7. The method according to claim 1, characterized in that It also includes a computational complexity optimization step: Based on the lightweight Transformer model structure, the computational complexity is controlled at O(n*d), where n is the length of the time series and d is the input feature dimension.
8. The method according to claim 1, characterized in that: The model optimization step also includes: The Transformer model is further optimized by combining local search strategy and gradient descent method; Use L1 regularization to prune unimportant model components and balance pruning and search tasks.
9. The method according to claim 1, characterized in that: It also includes the automatic learning strategy step: Automatically adjust the model search strategy based on the performance results of different training data; Dynamically update the search space and optimization objectives to adapt to different time series data characteristics.
10. The method according to claim 1, characterized in that The classification step further comprises: Perform real-time feature rearrangement on input time series data; Use the trained lightweight Transformer model to extract spatiotemporal features; Generate classification probability distribution through fully connected layers and softmax function; The category with the highest probability is selected as the final classification result.