Time sequence anomaly detection method and system based on combination of hierarchical adaptive attention and Mama
Through the method of combining hierarchical adaptive attention with Mamba, the problem of high computational complexity and lack of flexibility in fixed attention mechanism in time series anomaly detection is solved, efficient and accurate anomaly detection is achieved, and resource allocation can be dynamically adjusted on different time scales, capture long-term dependence and short-term fluctuations, and improve detection accuracy.
Patent Information
- Application Number
- CN202510547196.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-19
AI Technical Summary
The existing time series anomaly detection methods have high computational complexity when processing long sequences, the fixed attention mechanism lacks flexibility, it is difficult to adapt to the time characteristics of different data sets, and it is difficult to take into account both short-term fluctuations and long-term dependencies, resulting in a blind spot for detection.
Using a method of combining hierarchical adaptive attention with Mamba, the computing resources are dynamically allocated in the time context through a multi-grained token routing strategy, and combining self-attention and block-level selective attention mechanisms to build a time series anomaly detection model, including reconstruction error calculation, error normalization and hierarchical fraction fusion, adaptively adjust parameters to simulate long-distance dependencies and enhance nonlinear time mode modeling.
It improves the efficiency and accuracy of abnormal detection, can dynamically adjust focus on different time scales, effectively capture abnormal patterns in multivariable time series, reduces the computational complexity and improves the detection performance of subtle anomalies.
Smart Images

Figure CN120508850A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a time series anomaly detection technology in data mining, and specifically to a time series anomaly detection method and system combining hierarchical adaptive attention with Mamba. Background Art
[0002] Time series anomaly detection plays a vital role in industrial monitoring, healthcare, cybersecurity, and financial systems. In these fields, the ability to identify anomalous patterns serves multiple critical functions: preventing catastrophic system failures, detecting fraudulent activity, and maintaining operational efficiency. The task of time series anomaly detection goes beyond simple pattern recognition and has become a crucial component of modern system reliability and security frameworks.
[0003] Real-world temporal data poses significant challenges for continuously testing existing methods, primarily in three areas. First, nonstationary behavior—statistical properties that change unpredictably over time—produces moving detection targets that are difficult for fixed models to track. Second, subtle anomalies that deviate only slightly from normal patterns require increased sensitivity without increasing false positives. Third, multiscale temporal phenomena require simultaneous analysis across different time scales, from microsecond fluctuations to long-term trends.
[0004] Traditional statistical and machine learning techniques exhibit significant limitations when dealing with complex time series data. As demonstrated by traditional methods such as KNN-based methods and LOF, these techniques perform well on well-structured data but struggle to handle complex temporal dependencies in high-dimensional environments. Deep learning architectures have addressed some of these limitations, with recurrent neural network models such as LSTM-AD and MRA-LSTM improving temporal modeling, while OmniAnomaly enhances pattern recognition through random networks and normalized flows. However, these methods still face the challenge of scaling temporal dependencies. Transformer-based architectures such as the Anomaly Transformer and DCDetector have advanced the field by leveraging self-attention for temporal relationship modeling.
[0005] Although existing methods have achieved remarkable results, they still have three key problems:
[0006] (1) There is a common problem of quadratic computational complexity (O(n 2 )) problem. When processing longer time series, the computational cost of the model rises sharply, resulting in a significant decrease in training and inference efficiency.
[0007] (2) Using a fixed attention mechanism, such as local attention, sparse attention, or sliding window attention, although this preset attention structure can reduce the computational complexity to a certain extent, it lacks flexibility and is difficult to dynamically adjust the distribution of attention weights according to the temporal characteristics of different data sets, and cannot adapt to the different temporal characteristics between different data sets.
[0008] (3) It is difficult to simultaneously account for both short-term fluctuations and long-term dependencies in sequence modeling. This is particularly prominent in real-world time series applications, often leading to "detection blind spots" in the model at critical moments. Short-term fluctuations usually reflect sudden events or local anomalies, while long-term dependencies reveal trend changes, cyclical patterns, or delayed effects. Summary of the Invention
[0009] Purpose of the invention: To address the shortcomings of existing time series anomaly detection methods based on machine learning, reconstruction, and Transformer structure, the present invention provides a time series anomaly detection method and system based on hierarchical adaptive attention combined with Mamba.
[0010] Technical Solution: A time series anomaly detection method based on the combination of hierarchical adaptive attention and Mamba. The construction of the time series anomaly detection model includes the following steps:
[0011] S1, based on Mamba, captures long-term dependencies while enhancing the ability to model nonlinear temporal patterns. It has two parallel channels: Channel 1 extracts temporal features through one-dimensional convolution and linear projection, while Channel 2 dynamically generates the four key parameter matrices required to build the state space model based on the input content.
[0012] S2. Dynamically allocate computing resources in the temporal context based on a multi-granularity token routing strategy, and combine self-attention with a block-level selective attention mechanism to achieve comprehensive coverage of temporal information and improve computational efficiency.
[0013] The self-attention mechanism channel is calculated based on the original attention matrix and uses a causal mask to model standard temporal dependencies. The block-level selective attention mechanism channel selectively identifies the most relevant block information through a gating mechanism and performs attention calculations within these highly correlated regions without causal constraints, thereby capturing the complex dependencies between different temporal contexts.
[0014] S3. Build a time series anomaly detection model. This model includes reconstruction error calculation, error normalization, and anomaly score fusion for time series anomaly detection. Specifically:
[0015] (1) Reconstruction error calculation
[0016] Represents a multivariate time series, where L represents the sequence length and D represents the number of variables. The reconstructed representation is The reconstruction error is:
[0017]
[0018] (2) Error normalization is performed based on context awareness. This process considers the local feature statistics and global distribution properties of the time series and is expressed as:
[0019]
[0020] where μ d and σ d represents the global mean and standard deviation of dimension d in the entire sequence, μ t,w,d and σ t,w,d represents the local statistics calculated within a sliding window of size w centered at time t, with a parameter γ of dimension d ∈[0,1] balances global and local normalization according to the inherent characteristics of each dimension and learns to update during training. e(t,d) represents the reconstruction error at time t and dimension d.
[0021] (3) The anomaly score integrates information from multiple scales through a hierarchical fusion mechanism. The calculation of the anomaly score is:
[0022]
[0023] The above formula combines the Euclidean norm of all dimensions with the maximum normalized error, and the hyperparameter α∈[0,1] is used to control the relative contribution of each component in the anomaly score;
[0024] In order to further improve the detection performance of subtle anomalies, an attention weighting strategy based on the AnomalyAttention mechanism is introduced. The anomaly score weighting process is as follows:
[0025]
[0026] Among them, AttentionScore t Quantify the degree of attention deviation at time t compared to the sequence average, β is used as the influencing factor, and finally the threshold θ is applied to the final anomaly score for anomaly detection, which is expressed as:
[0027]
[0028] Furthermore, the time series anomaly detection model includes three parts: initial embedding, feature encoding, and anomaly score calculation:
[0029] Time series data is first transformed through a data embedding layer, which projects the original input features into a higher-dimensional representation suitable for subsequent processing;
[0030] The feature encoding consists of stacked encoder layers, with two parallel processing branches to process the Mamba state space model and the abnormal attention mechanism. The two are processed in parallel on the input representation, and then their outputs are dynamically fused through an adaptive gating mechanism. The calculation formula of this process is as follows:
[0031] x′=LayerNorm(MambaSSM(x)+AnomalyAttention(x))
[0032] Where x represents the layer input, x′ represents the output after the expansion layer processing, the outputs of the two branches are normalized by the layer, and then weighted fused through the adaptive gating mechanism to achieve dynamic regulation of the information contribution of each branch. The mathematical expression is:
[0033]
[0034] Where x represents the normalized output, x skip is the skip connection from the input, W g and b g is a learnable parameter and σ is an activation function. The encoder consists of N identical encoding layers stacked sequentially, with each layer refining the feature representation;
[0035] After the encoder, the feature representation goes through layer normalization, a feedforward network, and a projection layer to generate a reconstruction result. The final anomaly detection stage calculates the anomaly score by comparing the reconstructed output with the original input and exploiting the deviations at the sequence level and feature level.
[0036] Furthermore, the expanded feature representation is fed into the Mamba state-space model, which consists of two parallel channels:
[0037] Channel 1 realizes temporal feature extraction through one-dimensional convolution and linear projection;
[0038] Channel 2 dynamically generates the four key parameter matrices required to build the state space model based on the input content:
[0039] A t ,B t ,C t ,D t =ParameterGeneration(x expand )
[0040] The above parameters enable the model to adaptively adjust its dynamic behavior according to the input, thus enhancing the flexibility of temporal modeling;
[0041] The state update process follows the following state space equation:
[0042] h t =A t h t-1 +B t X t
[0043] y t =C t h t +D t x t
[0044] where h t represents the hidden state at time step t, y t To output the result;
[0045] Afterwards, the feature outputs generated by the two channels are fused through a selective scan operation. The processed feature representations are mapped back to the original dimensions through a projection layer and combined with the input through a residual connection to generate the final output of the Mamba state space model:
[0046] x mamba =x+ProjectBack(y)
[0047] Where ProjectBack(y) represents the linear projection operation.
[0048] Furthermore, the abnormal attention mechanism includes strategically combining self-attention with block-level selective attention mechanism by introducing a hierarchical and multi-granular token routing strategy to achieve comprehensive coverage of temporal information and computational efficiency. The specific calculation process formula is as follows:
[0049] O=AnomalyAttention(Q,K,V)=combine(O s ,O m )
[0050] Among them O s represents the self-attention output, O m Represents the block attention output, which is combined by adaptive weighting based on the input context;
[0051] In the block partitioning and selection strategy, the key-value pairs are divided into blocks B of different sizes, where the total number of blocks is At the same time, local queries are introduced to implement local fine-grained dependency modeling, and global queries are introduced to capture global context information. These two types of queries are dynamically routed by query routers. The time series anomaly detection model can adaptively balance local and global attention.
[0052] Furthermore, the method uses the following method to filter the blocks divided by the key value to obtain the key block:
[0053] The gating score is first calculated by average pooling the key representations: Then calculate the attention score between the query and the collection key: The causal mask ensures that the model does not pay attention to future blocks and selects the most relevant blocks G = topk(S+M,k) at each query position.
[0054] According to the above method, the present invention can also realize a time series anomaly detection system based on hierarchical adaptive attention combined with Mamba, including the following components:
[0055] Mamba state space module: dual-channel structure, performs step S1 of the method;
[0056] An abnormal attention mechanism module, executing step S2 of the method;
[0057] The anomaly score calculation module executes step S3 of the method.
[0058] Beneficial effects: Compared with the prior art, the substantial features and effects of the present invention are:
[0059] (1) Enhanced Mamba integration improves the standard Mamba so that its parameters can be dynamically adjusted according to the characteristics of the input sequence, effectively simulating long-distance dependencies while enhancing the ability to model nonlinear temporal patterns.
[0060] (2) Abnormal attention mechanism, which dynamically allocates computing resources in the temporal context by introducing a hierarchical, multi-granular token routing strategy, adaptively focusing processing power on key information fragments while retaining a broad global perception. This mechanism can dynamically adjust the focus of attention at different time scales and modes according to the complexity of the input data.
[0061] The present invention adopts a multi-level scoring mechanism, which comprehensively considers the temporal context and dataset characteristics. The anomaly score calculation process includes three stages: reconstruction error calculation, error normalization and hierarchical score fusion, which improves the efficiency and accuracy of anomaly detection and provides effective support for the development of time series anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 This is a schematic diagram of the overall process framework of the present invention;
[0063] Figure 2 Shown is a schematic diagram of the Mamba module in the present invention;
[0064] Figure 3Shown is a schematic diagram of the AnomalyAttention module in the present invention. DETAILED DESCRIPTION
[0065] This paper provides a time series anomaly detection method based on hierarchical adaptive attention combined with Mamba for robust anomaly detection. Generally speaking, this paper determines the anomaly score by constructing the following three modules:
[0066] Mamba State-Space Module (Mamba SSM): This paper proposes an enhanced Mamba integration module that effectively captures long-term dependencies while enhancing the ability to model nonlinear temporal patterns. The Mamba State-Space Module is capable of processing long sequences, avoiding the quadratic computational complexity of traditional Transformer architectures while maintaining high sensitivity to subtle pattern deviations that characterize anomalous real-world behavior.
[0067] Anomaly Attention: This paper proposes an anomaly attention mechanism that dynamically allocates computational resources within a temporal context by introducing a multi-granularity token routing strategy. This attention mechanism focuses computational resources on key pieces of information while maintaining global context awareness, achieving fine-grained attention allocation across multiple timescales.
[0068] Anomaly Score Calculation Module: This paper employs a multi-layered scoring mechanism, encompassing reconstruction error calculation, error normalization, and hierarchical score fusion. This mechanism comprehensively considers temporal context and dataset characteristics, effectively capturing anomalous patterns across all dimensions of multivariate time series.
[0069] The implementation process of the method of the present invention will be described in detail below with reference to embodiments.
[0070] (1) Problem definition
[0071] set up represents a multivariate time series, where L represents the sequence length and D represents the number of variables. The time series anomaly detection task aims to identify a binary label sequence y = {y1, y2, ..., y L}, where y t ∈{0,1} means that the observation at time t is abnormal (y t =1) or normal (y t =0). Given an input time series X, the present invention processes it through a series of transformations and calculates anomaly scores s = {s1, s2, ..., s L Each score is compared with an optimal threshold. When the score exceeds this threshold, the corresponding time point is classified as an anomaly; otherwise, it is considered normal behavior.
[0072] (2) Overall framework of the method
[0073] The present invention consists of three main parts: initial embedding, feature encoding with dedicated layers, and final anomaly score calculation. The input features are first transformed through a data embedding layer, where L represents the sequence length and D represents the number of variables. This layer projects the raw input features into a higher-dimensional representation suitable for subsequent processing. The embedded representation then passes through an encoder consisting of stacked encoder layers. Each encoder layer forms the core of the architecture and contains two parallel processing branches: a Mamba SSM block and an anomaly attention layer. These components operate simultaneously on the input representation and are subsequently combined through an adaptive gating mechanism. The formula is as follows:
[0074] x′=LayerNorm(MambaSSM(x)+AnomalyAttention(x))
[0075] Where x represents the layer input and x′ represents the processed output. The outputs of the two branches are normalized by the layer, and then the contributions of the two are dynamically balanced through an adaptive gating mechanism:
[0076]
[0077] Where x represents the normalized output, x skip is the skip connection from the input, W g and b g is a learnable parameter, and σ is the activation function. The encoder consists of N identical encoding layers stacked sequentially, each of which refines the feature representation. After the encoder, the representation undergoes layer normalization, a feedforward network, and a projection layer to produce a reconstruction. The final anomaly detection stage compares the reconstructed output with the original input, leveraging both sequence-level and feature-level deviations to compute an anomaly score.
[0078] (3) Mamba SSM module
[0079] Specifically, the Mamba State Space Model (SSM) block is a fundamental component in our architecture that enables efficient modeling of long-term dependencies in time series data. The Mamba SSM block implements a selective state space model that adaptively processes sequential information. The module first transforms the input through an expansion layer that increases representational power:
[0080] x expanded =Expand(x)
[0081] The expanded representation then passes through a state-space model (SSM), which consists of parallel processing paths. One path applies a one-dimensional convolution followed by a linear projection, while the other generates content-dependent parameters for the state update mechanism.
[0082] The parameter generation path produces four key parameter matrices that define the state-space dynamics:
[0083] A t ,B t ,C t ,D t =ParameterGeneration(x expand )
[0084] These content-dependent parameters enable the model to adapt its dynamics based on the input, allowing for more flexible and expressive temporal modeling. The state update mechanism implements the following state-space equations:
[0085] h t =A t h t-1 +B t X t
[0086] y t =C t h t +D t x t
[0087] where h t represents the hidden state at time step t, y t The selective scan operation efficiently implements these recursive equations using parallelizable operations, significantly reducing the computational complexity compared to traditional recursive architectures.
[0088] The outputs from the two processing paths are combined through a selective scan operation, which maintains the sequential information flow while enabling efficient parallel computation. Finally, the processed representation is projected to the original dimension through a projection layer and combined with the input through a residual connection to produce the Mamba output:
[0089] x mamba =x+ProjectBack(y)
[0090] This state-space formulation offers several advantages for time series modeling: (1) it effectively captures long-range dependencies without the vanishing gradient problem of traditional RNNs; (2) content-dependent parameters allow adaptive processing based on the input pattern; and (3) the selective scanning operation enables efficient parallel computation, making the model scalable to long sequences.
[0091] (4) Anomaly Attention Module
[0092] The anomaly attention layer captures irregular patterns in multivariate time series through an effective block attention method that focuses on relevant temporal blocks while keeping the computation tractable.
[0093] This component handles the hidden state of the input Where N represents the sequence length, h represents the number of attention heads, and d represents the head size. Unlike the standard self-attention calculation at all positions, this method strategically combines full self-attention and selective block-based attention to achieve comprehensive coverage and computational efficiency. The specific formula is formalized as:
[0094] O=AnomalyAttention(Q,K,V)=combine(O s ,O m )
[0095] Among them O s represents the standard self-attention output, O m Represents the block attention output, which is combined by adaptive weighting based on the input context. One of the notable features is the block division and selection strategy. The key-value pairs are divided into blocks of different sizes B, where the total number of blocks is Meanwhile, the query representation process through local queries (Q1) focuses on fine-grained local dependencies, and global queries aims to model broader contextual patterns (Q2). These query representations are dynamically routed through the query router module, enabling the model to adaptively balance local and global attention patterns.
[0096] To efficiently select a specific block, the gating score is first calculated by applying average pooling to the key representation: Then calculate the attention score between the query and the collection key: A causal mask is used to ensure that future blocks are not attended to, and the most relevant blocks G = topk(S+M,k) are selected at each query position.
[0097] The anomaly attention layer implements two parallel attention channels. The self-attention channel performs standard self-attention using the original matrix that preserves causal constraints. Meanwhile, the block attention channel selectively processes only the most relevant blocks identified by the gating mechanism, without causal constraints in these highly relevant segments. This design is able to capture complex dependencies between different temporal contexts.
[0098] The standard self-attention component ensures comprehensive sequence modeling, while the block attention mechanism enables efficient processing of long sequences by focusing computational resources on relevant time segments. The dual-channel design simultaneously captures fine-grained sequence patterns and broader contextual relationships, which is crucial for identifying different anomaly types. In addition, the top-k block selection method significantly reduces computational complexity, making the processing of long time series more efficient.
[0099] The anomaly attention mechanism effectively combines the Mamba SSM module. The Mamba module captures the sequence dynamics through its state-space formulation, while the anomaly attention layer specifically identifies contextual deviations and irregular patterns that are crucial for anomaly detection, and is used to model normal behaviors and anomalies in multivariate time series.
[0100] (5) Anomaly score calculation
[0101] This paper adopts a multi-level scoring mechanism to effectively capture anomalies in all dimensions of multivariate time series. The anomaly detection process consists of three stages: reconstruction error calculation, error normalization, and hierarchical score fusion.
[0102] Original input time series and the reconstruction output of the model First calculate the point-by-point reconstruction error:
[0103]
[0104] Instead of simply using the resulting raw error, our approach implements a context-aware normalization process that takes into account both local and global statistical properties:
[0105]
[0106] where μ d and σ d represents the global mean and standard deviation of dimension d in the entire sequence, and μ t,w,d and σ t,w,d represents the local statistics computed within a sliding window of size w centered at time t. The dimension-specific parameters γ d ∈[0,1] balances global and local normalization according to the inherent characteristics of each dimension and is learned to update during training.
[0107] The final anomaly score integrates information from multiple scales through a hierarchical fusion mechanism:
[0108]
[0109] This formula combines the Euclidean norm of all dimensions with the maximum normalized error, and the hyperparameter α∈[0,1] controls the relative contribution of each component. In order to further improve the detection performance of subtle anomalies, an attention weighting based on the AnomalyAttention mechanism is introduced:
[0110]
[0111] Among them, AttentionScore t quantifies the degree of attention deviation at time t compared to the sequence mean, with β being the influencing factor. This attention-enhanced score amplifies the detection signal of points that receive anomalous attention from the model, effectively leveraging the learned representation to identify subtle pattern deviations that reconstruction error alone might miss. A threshold θ is then applied to the final anomaly score to perform anomaly detection:
[0112]
[0113] The present invention implements an adaptive threshold selection strategy that takes into account temporal context and dataset characteristics, optimizing the precision-recall trade-off for each specific application domain.
[0114] Example
[0115] Datasets: To evaluate the performance of the proposed model, this example is evaluated on five widely used multivariate time series anomaly detection benchmark datasets: SMD (Server Machine Dataset), MSL (Mars Science Laboratory), SMAP (Soil Moisture Active Passive Satellite), SWaT (Safe Water Treatment), and PSM (Aggregated Server Metrics). These datasets span diverse domains, including industrial systems, spacecraft telemetry, and server monitoring, and have different characteristics in terms of dimensionality, temporal patterns, and anomaly types.
[0116] Comparison Models: This example compares the proposed model with current SOTA models, including BOCPD, LSTM-VAE, BeatGAN, LSTM, OmniAnomaly, InterFusion, THOC, AnomalyTrans, DCdetector, and MAAT, a total of 10 methods, covering various methods based on distance, density, reconstruction, and contrastive learning in time series anomaly detection.
[0117] Evaluation Metrics: This example uses precision (P), recall (R), and F1 score (F1) as evaluation metrics. Precision measures the proportion of correctly identified anomalies among all detected anomalies, while recall measures the percentage of correctly identified anomalies among all actual anomalies. The F1 score provides a balanced measure of precision and recall.
[0118] The proposed method excels across five benchmark datasets, consistently outperforming baseline methods in terms of precision, recall, and F1 score. It achieves high recall rates of 99.18% and 100% on the PSM and SWaT datasets, respectively, highlighting its superior ability to detect anomalies in mission-critical scenarios. Compared to other baselines, including traditional machine learning, reconstruction-based, and contrastive learning methods, the proposed method demonstrates consistent accuracy advantages and effectively reduces false negatives. Combining the sequential modeling capabilities of the Mamba state-space model with the pattern recognition strength of the Anomaly Attention mechanism, the proposed method improves computational efficiency while focusing attention on critical time periods. Experimental results confirm that the proposed method provides robust and comprehensive performance across a variety of complex anomaly detection tasks.
[0119] Table 1 Performance comparison of different models on different datasets
[0120]
Claims
1. A time series anomaly detection method based on hierarchical adaptive attention combined with Mamba, characterized by: The construction of a time series anomaly detection model includes the following steps: S1, based on Mamba, captures long-term dependencies while enhancing the ability to model nonlinear temporal patterns. It has two parallel channels: Channel 1 extracts temporal features through one-dimensional convolution and linear projection, while Channel 2 dynamically generates the four key parameter matrices required to build the state space model based on the input content. S2. Dynamically allocate computing resources in the temporal context based on a multi-granularity token routing strategy, and combine self-attention with a block-level selective attention mechanism to achieve comprehensive coverage of temporal information and improve computational efficiency. The self-attention mechanism channel is calculated based on the original attention matrix and uses a causal mask to model standard temporal dependencies. The block-level selective attention mechanism channel selectively identifies the most relevant block information through a gating mechanism and performs attention calculations within these highly correlated regions without causal constraints, thereby capturing the complex dependencies between different temporal contexts. S3. Build a time series anomaly detection model. This model includes reconstruction error calculation, error normalization, and anomaly score fusion for time series anomaly detection. Specifically: (1) Reconstruction error calculation Represents a multivariate time series, where L represents the sequence length and D represents the number of variables. The reconstructed representation is The reconstruction error is: (2) Error normalization is performed based on context awareness. This process considers the local feature statistics and global distribution properties of the time series and is expressed as: where μ d and σ d represents the global mean and standard deviation of dimension d in the entire sequence, μ t,w,d and σ t,w,d represents the local statistics calculated within a sliding window of size w centered at time t, with a parameter γ of dimension d ∈[0,1] balances global and local normalization according to the inherent characteristics of each dimension and learns to update during training. E(t,d) represents the reconstruction error at time t and dimension d; (3) The anomaly score integrates information from multiple scales through a hierarchical fusion mechanism. The calculation of the anomaly score is: The above formula combines the Euclidean norm of all dimensions with the maximum normalized error, and the hyperparameter α∈[0,1] is used to control the relative contribution of each component in the anomaly score; In order to further improve the detection performance of subtle anomalies, an attention weighting strategy based on the AnomalyAttention mechanism is introduced. The anomaly score weighting process is as follows: Among them, AttentionScore t Quantify the degree of attention deviation at time t compared to the sequence average, β is used as the influencing factor, and finally the threshold θ is applied to the final anomaly score for anomaly detection, which is expressed as:
2. The time series anomaly detection method according to claim 1, characterized in that: The time series anomaly detection model described consists of three parts: initial embedding, feature encoding, and anomaly score calculation: Time series data is first transformed through a data embedding layer, which projects the original input features into a higher-dimensional representation suitable for subsequent processing; The feature encoding consists of stacked encoder layers, with two parallel processing branches to process the Mamba state space model and the abnormal attention mechanism. The two are processed in parallel on the input representation, and then their outputs are dynamically fused through an adaptive gating mechanism. The calculation formula of this process is as follows: x′=LayerNorm(MambaSSM(x)+AnomalyAttention(x)) Where x represents the layer input, x′ represents the output after the expansion layer processing, the outputs of the two branches are normalized by the layer, and then weighted fused through the adaptive gating mechanism to achieve dynamic regulation of the information contribution of each branch. The mathematical expression is: Where x represents the normalized output, x skip is the skip connection from the input, W g and b g is a learnable parameter, σ is an activation function; The encoder consists of N identical encoding layers stacked sequentially. This step refines the feature representation for each layer. After the encoder, the feature representation undergoes layer normalization, a feedforward network, and a projection layer to generate a reconstruction result. The final anomaly detection stage calculates the anomaly score by comparing the reconstructed output with the original input and using the deviation between the sequence level and the feature level.
3. The time series anomaly detection method according to claim 1 or 2, characterized in that: The expanded feature representation is input into the Mamba state space model, which consists of two parallel channels: Channel 1 realizes temporal feature extraction through one-dimensional convolution and linear projection; Channel 2 dynamically generates the four key parameter matrices required to build the state space model based on the input content: A t ,B t ,C t ,D t =ParameterGeneration(x expand ) The above parameters enable the model to adaptively adjust its dynamic behavior according to the input, thus enhancing the flexibility of temporal modeling; The state update process follows the following state space equation: h t =A t h t-1 +B t X t y t =C t h t +D t x t where h t represents the hidden state at time step t, y t To output the result; Afterwards, the feature outputs generated by the two channels are fused through a selective scan operation. The processed feature representations are mapped back to the original dimensions through a projection layer and combined with the input through a residual connection to generate the final output of the Mamba state space model: x mamba =x+ProjectBack(y) Where ProjectBack(y) represents the linear projection operation.
4. The time series anomaly detection method according to claim 1, characterized in that: The abnormal attention mechanism includes introducing a hierarchical and multi-granular token routing strategy, strategically combining self-attention with a block-level selective attention mechanism to achieve comprehensive coverage of temporal information and computational efficiency. The specific calculation process formula is as follows: O=AnomalyAttention(Q,K,V)=combine(O s ,O m ) Among them O s represents the self-attention output, O m Represents the block attention output, which is combined by adaptive weighting based on the input context; In the block partitioning and selection strategy, the key-value pairs are divided into blocks B of different sizes, where the total number of blocks is At the same time, local queries are introduced to implement local fine-grained dependency modeling, and global queries are introduced to capture global context information. These two types of queries are dynamically routed by query routers. The time series anomaly detection model can adaptively balance local and global attention.
5. The time series anomaly detection method according to claim 4, characterized in that: This method uses the following method to filter the blocks divided by key values to obtain key blocks: The gating score is first calculated by average pooling the key representations: Then calculate the attention score between the query and the collection key: The causal mask ensures that the model does not pay attention to future blocks and selects the most relevant blocks G = topk(S+M,k) at each query position.
6. A time series anomaly detection system based on hierarchical adaptive attention combined with Mamba, characterized by: The system performs the method according to any one of claims 1 to 5, and includes the following components: Mamba state space module: dual-channel structure, executing step S1 of the method according to any one of claims 1 to 5; An abnormal attention mechanism module, performing step S2 of the method according to any one of claims 1 to 5; The anomaly score calculation module executes step S3 of the method according to any one of claims 1 to 5.
Citation Information
Cited By
Electrolytic bath health state online prediction method and system based on state space model
CN121167243A
Lithium battery SOH estimation method based on multi-scale dual time sequence Mama
CN121878509A
Lithium battery soh estimation method based on multi-scale dual timing mamba
CN121878509B
Wafer process fault early warning method and system based on mixed sequence decomposition and Mama architecture
CN122112929A
Weak supervision anomaly detection method based on double-branch feature fusion
CN122241609A