Multivariate time sequence adaptive anomaly detection system and method based on Mixer-Transform
By combining the Mixer and Transformer models in multivariate time series anomaly detection, capturing global dependencies and calculating adaptive thresholds, the problem of difficulty in considering time dependencies and variable interactions in the prior art is solved, and more accurate and flexible anomaly detection is achieved.
Patent Information
- Application Number
- CN202510142614.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
When detecting exceptions in multivariable time series, it is difficult to consider the complex interaction between time dependencies and variables at the same time, and traditional fixed threshold methods are difficult to adapt to the dynamic environment, resulting in false positives or missed reports.
A multivariate time series adaptive anomaly detection system based on Mixer-Transformer is adopted to alternately model the channel and time dimension information through the Mixer module, capture global dependencies, and calculate correlation differences in combination with the Transformer model to calculate the abnormal score. At the same time, an adaptive threshold update mechanism is adopted to adjust the abnormal judgment criteria according to dynamic changes in the data.
Effectively capture complex dynamic spatiotemporal features, enhance the accuracy and flexibility of abnormal detection, reduce false alarms and missed alarms, and adapt to changes in the dynamic environment.
Smart Images

Figure CN120067938A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multivariate time series anomaly detection, and relates to a multivariate time series adaptive anomaly detection system and method based on Mixer-Transformer. Background Art
[0002] With the rapid development of sensor technology and data acquisition capabilities, many critical systems (such as industrial production, intelligent transportation, financial transactions, and medical monitoring) are generating massive amounts of time series data. These data not only reflect the operating status of the system but also contain important information about device performance and potential faults. Detecting anomalies in these time series data is an important means to ensure the safe operation of the system, optimize performance, and avoid losses. To reduce false alarms in time series detection and accurately identify the source of anomalies, the use of artificial intelligence technology has become a promising solution. These technologies can analyze the strong temporal correlations and dynamic changes between different data frames. Existing detection methods face two main challenges, leading to false alarms or missed detections: (i) detecting anomalies in multivariate time series requires considering both temporal dependencies and complex interactions between variables simultaneously; (ii) traditional fixed-threshold methods often struggle to adapt to dynamic environments. Summary of the Invention
[0003] To solve these problems, the present invention proposes a multivariate time series adaptive anomaly detection system and method based on Mixer-Transformer, which combines the Mixer structure and the Transformer model. By alternately modeling the information interaction in the channel and time dimensions, it effectively captures global dependencies and enhances the model's ability to extract complex dynamic spatio-temporal features. At the same time, an adaptive threshold update mechanism is adopted to flexibly adjust the anomaly determination criteria according to the dynamic changes of the data.
[0004] The technical solution of the present invention is as follows:
[0005] A multivariate time series adaptive anomaly detection system based on Mixer-Transformer, characterized in that the system includes a Mixer module, a Transformer module, an anomaly score calculation module, and an adaptive threshold module;
[0006] The Mixer module is used to process the initial time series data. By alternately modeling the information interaction in the channel and time dimensions, it captures global dependencies and enhances the model's ability to extract complex dynamic spatio-temporal features;
[0007] The Transformer module is connected after the Mixer module and is used to calculate the correlation difference of the time series data processed by the Mixer module;
[0008] The abnormal score calculation module is used to calculate the abnormal score of time series data according to the correlation difference output by the Transformer module;
[0009] The adaptive threshold module is used to dynamically adjust the abnormal determination criterion according to the abnormal score.
[0010] The present invention also discloses a multi-variable time series adaptive anomaly detection method based on Mixer-Transformer, which uses the above detection system and includes the following steps:
[0011] (1), Use the Mixer module to process the initial time series data, and capture the global dependence by alternately modeling the interaction of channel and time dimension information, and output the processed time series data;
[0012] (2), Input the time series data processed by the Mixer module into the Transformer module, calculate the correlation difference of the time series data, where the calculation of the correlation difference includes comparing the difference between the prior correlation and the series correlation at each time point, and using the symmetric KL divergence to calculate the difference between the two;
[0013] (3), Input the correlation difference output by the Transformer module into the abnormal score calculation module to calculate the abnormal score of each time point in the time series data;
[0014] (4), Use the adaptive threshold module to dynamically adjust the abnormal determination criterion according to the abnormal score. The adaptive threshold module introduces the setting of a sliding window, uses the exponential weighted moving average method to smooth the abnormal score, and sets the response boundary according to the difference between the smoothed abnormal score and the moving average value to determine the abnormality.
[0015] Preferably, the specific steps of the above step (1) are as follows:
[0016] (1-1), The Mixer module processes the initial time series data. First, the input initial time series data to be determined is where B is the batch size, T is the number of time steps, and C is the number of feature channels;
[0017] (1-2), The time mixing sub-layer reshapes the input into the form of , and then passes through a fully connected network to obtain the output where t represents the t-th moment in the time series data, and l represents the number of layers of the Mixer;
[0018] (1-3), The calculation formula of is as follows:
[0019]
[0020] Among them and are the weight matrices of the fully connected layers, D h is the hidden layer dimension, and σ is the activation function;
[0021] (1-4) The operation of the feature mixing sublayer is similar to that of the temporal mixing sublayer, except that the input features are reshaped into form, and then the output
[0022] (1-5) The calculation formula of is as follows:
[0023]
[0024] Among them and are the weight matrices of the fully connected layers, D h is the hidden layer dimension, and σ is the activation function;
[0025] (1-6) The output of the l-th layer Mixer is obtained through residual processing and normalization
[0026] (1-7)、 The calculation formula of is as follows:
[0027]
[0028] where LayerNorm() represents the normalization operation, and Reshape() reshapes the tensor into form;
[0029] (1-8) Placing N stacked Mixers in front of the Transformer module constitutes a preprocessing module for time series data, denoted as MixerPretrain;
[0030] (1-9) The calculation formula of MixerPretrain is as follows:
[0031]
[0032] Among them is the original multivariate time series data, C is the original feature dimension, and l ∈ (1, 2,..., N); Time series features containing more global context information are used to replace the original multivariate time series data and input into the Transformer module.
[0033] Preferably, the specific steps of the above step (2) are as follows:
[0034] (2-1) Calculate the association difference at each time point in the time series by comparing the differences between the prior association and the series association at each time point, and use the symmetric KL divergence to calculate the difference between the two;
[0035] (2-2) The prior association reflects the expected association pattern between time points in the time series within a local range. Calculate the prior association of each time point relative to other time points through a learnable Gaussian kernel;
[0036]
[0037] where i, j represent the i-th time point and the j-th time point, and σ i is the learnable scale parameter corresponding to the i-th time point. Based on the characteristics of the Gaussian distribution, calculate the association weight according to the relative distance |j - i| between time points. The closer the distance, the greater the weight, indicating that the association between adjacent time points in the time series is closer. Rescale[] is the form of converting the association weight matrix into a discrete distribution;
[0038] (2-3) The series association is the association relationship adaptively learned from the original time series data, which can capture the dynamic dependence relationship in the time series and reflect the true degree of association between time points; the calculation process of the series association is implemented through a standard self-attention mechanism; its calculation formula is:
[0039]
[0040] where represents the output of the (l - 1)-th layer, l ∈ {1, 2, …, L}, is the learnable parameter matrix, and the dimensions of Q, K, and V are N * d model , N is the length of the time series, and d model is the number of channels of the hidden state in the model. QK T represents the dot product of the query Q and the key K, obtaining an attention score matrix with dimensions N * N;
[0041] (2-4) In the multi-layer Mixer structure, calculate the prior association PA and the series association SA for each layer; PA and SA are the sets of calculation results for each layer, expressed as:
[0042] PA = {PA 1 , PA 2 , …, PA L}
[0043] SA = {SA 1 , SA 2 , …, SA L}
[0044] (2-5), the symmetric KL divergence is used to calculate the difference between the prior correlation and the serial correlation, while considering and The KL divergence at the i-th time point in the l-th layer is defined as:
[0045]
[0046] (2-6), considering the correlation information learned at different levels in the multi-layer model, the correlation differences of each layer are fused to obtain a more representative and stable correlation difference metric where L represents the number of layers of the multi-layer model;
[0047]
[0048] Preferably, in the anomaly score calculation module in step (3) above, the anomaly score of each time point in the time series is calculated from the correlation difference; the anomaly score of each time point in the time series is defined as:
[0049]
[0050] where is the reconstruction result of the input time series data .
[0051] 6. The multivariate time series adaptive anomaly detection method based on Mixer-Transformer according to claim 5, wherein the specific steps of step (4) are as follows:
[0052] (4-1), introduce the setting of a sliding window. At each moment t, observe the anomaly score sequence {S t-w , …, S t , … S t+w} within the window of length w before and after it;
[0053] (5-2), adopt the exponentially weighted moving average method to smooth the input values in order to better capture the long-term trend in the data; given the anomaly score S t at the current time step, its corresponding exponentially weighted moving average value EWMA t can be updated through the following recursive relationship:
[0054] EWMA t = αEWMA t-1 + (1 - α)S t
[0055] where EWMA t-1is the exponentially weighted moving average at the previous moment, α ∈ (0, 1) is the smoothing factor. A larger α value means the model pays more attention to the recent data points, while a smaller α value increases the dependence on historical data, making the model smoother;
[0056] (5 - 3) Set a sensitive response boundary threshold based on the difference between the current input value and the moving average. The update formula for the threshold is as follows:
[0057]
[0058] where δ t is the threshold at the current moment, used to determine whether the current data significantly deviates from the expected trend. Calculate the current anomaly score S t and the absolute deviation between the moving average EWMA t , that is, the distance between the current observation value and the smoothed trend.
[0059] Advantages of the present invention:
[0060] The time - series anomaly detection method based on Mixer - Transformer of the present invention combines the Mixer structure and the Transformer model. By alternately modeling the information interaction of the channel and time dimensions, it effectively captures the global dependence and enhances the model's ability to extract complex dynamic spatio - temporal features. At the same time, an adaptive threshold update mechanism is adopted to flexibly adjust the anomaly determination criterion according to the dynamic changes of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is the execution flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0062] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.
[0063] A multivariate time - series adaptive anomaly detection system based on Mixer - Transformer, characterized in that the system includes a Mixer module, a Transformer module, an anomaly score calculation module, and an adaptive threshold module;
[0064] The Mixer module is used to process the initial time - series data. By alternately modeling the information interaction of the channel and time dimensions, it captures the global dependence and enhances the model's ability to extract complex dynamic spatio - temporal features;
[0065] The Transformer module, connected after the Mixer module, is used to calculate the correlation difference of the sequential data processed by the Mixer module;
[0066] The anomaly score calculation module is used to calculate the anomaly score of the sequential data according to the correlation difference output by the Transformer module;
[0067] The adaptive threshold module is used to dynamically adjust the anomaly determination criterion according to the anomaly score.
[0068] The present invention also discloses a multivariate time series adaptive anomaly detection method based on Mixer-Transformer, which uses the above detection system, as Figure 1 shown, and includes the following steps:
[0069] (1), Use the Mixer module to process the initial sequential data, and by alternately modeling the interaction of channel and time dimension information, capture the global dependence, and output the processed sequential data;
[0070] (2), Input the sequential data processed by the Mixer module into the Transformer module to calculate the correlation difference of the sequential data, where the calculation of the correlation difference includes comparing the difference between the prior correlation and the series correlation at each time point, and using the symmetric KL divergence to calculate the difference between the two;
[0071] (3), Input the correlation difference output by the Transformer module into the anomaly score calculation module to calculate the anomaly score of each time point in the sequential data;
[0072] (4), Use the adaptive threshold module to dynamically adjust the anomaly determination criterion according to the anomaly score, where the adaptive threshold module introduces the setting of a sliding window, uses the exponential weighted moving average method to smooth the anomaly score, and sets the response limit according to the difference between the smoothed anomaly score and the moving average value to determine the anomaly.
[0073] Preferably, the specific steps of the above step (1) are as follows:
[0074] (1-1), The Mixer module processes the initial sequential data. First, the input initial sequential data to be determined is where B is the batch size, T is the number of time steps, and C is the number of feature channels;
[0075] (1-2), The time mixing sublayer reshapes the input into the form of , and then passes through a fully connected network to obtain the output where t represents the t-th moment in the sequential data, and l represents the number of layers of the Mixer;
[0076] (1 - 3), The calculation formula of
[0077]
[0078] where and are the weight matrices of the fully - connected layer, D h is the hidden layer dimension, and σ is the activation function;
[0079] (1 - 4) The operation of the feature mixing sub - layer is similar to that of the time mixing sub - layer, except that the input features are reshaped into form, and then the output
[0080] (1 - 5) The calculation formula of
[0081]
[0082] where and are the weight matrices of the fully - connected layer, D h is the hidden layer dimension, and σ is the activation function;
[0083] (1 - 6) The output of the l - th layer Mixer is obtained through residual processing and normalization
[0084] (1 - 7), The calculation formula of
[0085]
[0086] where LayerNorm() represents the normalization operation, and Reshape() reshapes the tensor into form;
[0087] (1 - 8) Placing N stacked Mixers in front of the Transformer module constitutes a pre - processing module for time - series data, denoted as MixerPretrain;
[0088] (1 - 9) The calculation formula of MixerPretrain is as follows:
[0089]
[0090] where is the original multi - variable time - series data, C is the original feature dimension, and l ∈ (1, 2, …, N); Temporal features containing more global context information replace the original multivariate time series data as the input to the Transformer module.
[0091] Preferably, the specific steps of the above step (2) are as follows:
[0092] (2-1) Calculate the association difference at each time point in the time series by comparing the differences between the prior association and the series association at each time point, and calculate the difference between the two using symmetric KL divergence;
[0093] (2-2) The prior association reflects the expected association pattern within a local range between time points in the time series. Calculate the prior association of each time point relative to other time points through a learnable Gaussian kernel;
[0094]
[0095] where i, j represent the i-th time point and the j-th time point, and σ i is the learnable scale parameter corresponding to the i-th time point. Based on the characteristics of the Gaussian distribution, calculate the association weight according to the relative distance |j - i| between time points. The closer the distance, the greater the weight, indicating that the association between adjacent time points in the time series is closer. Rescale[] is the form of converting the association weight matrix into a discrete distribution;
[0096] (2-3) The series association is an association relationship adaptively learned from the original time series data, which can capture the dynamic dependence relationship in the time series and reflects the true degree of association between time points; the series association calculation process is implemented through a standard self-attention mechanism; its calculation formula is:
[0097]
[0098] where represents the output of the (l - 1)-th layer, l ∈ {1, 2, …, L}, is the learnable parameter matrix, and the dimensions of Q, K, and V are N * d model , N is the length of the time series, and d model is the number of channels of the hidden state in the model. QK T represents the dot product of the query Q and the key K, resulting in an attention score matrix with dimensions N * N;
[0099] (2-4) In the multi-layer Mixer structure, calculate the prior association PA and the series association SA for each layer; PA and SA are the sets of calculation results for each layer, expressed as:
[0100] PA = {PA 1 , PA 2 , …, PA L}
[0101] SA = {SA 1 , SA 2 , …, SA L}}
[0102] (2 - 5), Use symmetric KL divergence to calculate the difference between the prior association and the serial association, while considering and The KL divergence at the i-th time point in the l-th layer is defined as:
[0103]
[0104] (2 - 6), Consider the association information learned at different levels in the multi-layer model, and fuse the association differences of each layer to obtain a more representative and stable association difference measure where L represents the number of layers of the multi-layer model;
[0105]
[0106] Preferably, in the anomaly score calculation module in step (3) above, the anomaly score of each time point in the time series is calculated from the association difference; the anomaly score of each time point in the time series is defined as:
[0107]
[0108] where is the reconstruction result of the input time series data .
[0109] 6. According to the multi-variate time series adaptive anomaly detection method based on Mixer-Transformer described in claim 5, the specific steps of step (4) are as follows:
[0110] (4 - 1), Introduce the setting of a sliding window. At each moment t, observe the anomaly score sequence {S t-w , …, S t , … S t+w} within the window of length w before and after it;
[0111] (5 - 2), Use the exponentially weighted moving average method to smooth the input values in order to better capture the long-term trend in the data; given the anomaly score S t at the current time step, its corresponding exponentially weighted moving average value EWMA t can be updated through the following recursive relationship:
[0112] EWMA t = αEWMA t-1+(1-α)S t
[0113] where EWMA t-1 is the exponentially weighted moving average at the previous moment, α ∈ (0, 1) is the smoothing factor. A larger α value means the model pays more attention to the recent data points, while a smaller α value increases the dependence on historical data, making the model smoother;
[0114] (5 - 3) Set a sensitive response boundary threshold based on the difference between the current input value and the moving average. The update formula for the threshold is as follows:
[0115]
[0116] where δ t is the threshold at the current moment, used to determine whether the current data significantly deviates from the expected trend. Calculate the absolute deviation between the current anomaly score S t and the moving average EWMA t That is, the distance between the current observed value and the smoothed trend.
[0117] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A multivariate time series adaptive anomaly detection system based on Mixer-Transformer, characterized by The system includes a Mixer module, a Transformer module, an anomaly score calculation module, and an adaptive threshold module; The Mixer module is used to process the initial time series data, capture global dependencies by alternately modeling channel and time dimension information interaction, and enhance the model's ability to extract complex dynamic spatiotemporal features; The Transformer module is connected to the Mixer module and is used to calculate the correlation difference of the time series data after being processed by the Mixer module; The anomaly score calculation module is used to calculate the anomaly score of the time series data according to the correlation difference output by the Transformer module; The adaptive threshold module is used to dynamically adjust the anomaly determination standard according to the anomaly score.
2. A multivariate time series adaptive anomaly detection method based on Mixer-Transformer, characterized in that: The detection system according to claim 1 comprises the following steps: (1) Use the Mixer module to process the initial time series data, capture global dependencies by alternately modeling channel and time dimension information interactions, and output the processed time series data; (2) Input the time series data processed by the Mixer module into the Transformer module to calculate the correlation difference of the time series data. The calculation of the correlation difference includes comparing the difference between the prior correlation and the series correlation at each time point, and using the symmetric KL divergence to calculate the difference between the two. (3) Input the correlation difference output by the Transformer module into the anomaly score calculation module to calculate the anomaly score at each time point in the time series data; (4) Use an adaptive threshold module to dynamically adjust the anomaly judgment standard according to the anomaly score. The adaptive threshold module introduces the setting of a sliding window, uses the exponentially weighted moving average method to smooth the anomaly score, and sets the response limit based on the difference between the smoothed anomaly score and the moving average to judge the anomaly.
3. The method for adaptive anomaly detection of multivariate time series based on Mixer-Transformer according to claim 2, characterized in that: The specific steps of step (1) are as follows: (1-1) The Mixer module processes the initial time series data. First, the input initial time series data with judgment is Where B is the batch size, T is the number of time steps, and C is the number of feature channels; (1-2), the time mixing sublayer reshapes the input into In the form of, and then through a layer of fully connected network to get the output Where t represents the time t in the time series data, and l represents the number of layers of Mixer; (1-3), The calculation formula is as follows: in and is the weight matrix of the fully connected layer, D h is the hidden layer dimension, σ is the activation function; (1-4) The operation of the feature mixing sublayer is similar to that of the time mixing sublayer, except that the input features are reshaped into In the form of, and then through two layers of fully connected network to get the output (1-5) The calculation formula is as follows: in and is the weight matrix of the fully connected layer, D h is the hidden layer dimension, σ is the activation function; (1-6) The output of the l-th layer Mixer is obtained through residual processing and normalization (1-7), The calculation formula is as follows: LayerNorm() represents the normalization operation, and Reshape() reshapes the tensor into form; (1-8) Place N stacked Mixers before the Transformer module to form a time series data preprocessing module, denoted as MixerPretrain; (1-9), the calculation formula of MixerPretrain is as follows: in is the original multivariate time series data, C is the original feature dimension, l∈(1, 2, …, N); The time series features containing more global context information replace the original multivariate time series data input into the Transformer module.
4. The method for adaptive anomaly detection of multivariate time series based on Mixer-Transformer according to claim 3, characterized in that: The specific steps of step (2) are as follows: (2-1) Calculate the correlation difference at each time point in the time series by comparing the difference between the prior correlation and the series correlation at each time point, and use the symmetric KL divergence to calculate the difference between the two; (2-2) Prior correlation reflects the expected correlation pattern between time points in the time series in a local range, and the prior correlation of each time point relative to other time points is calculated through a learnable Gaussian kernel; Where i, j represents the i-th time point and the j-th time point, σ i is the learnable scale parameter corresponding to the i-th time point. Based on the characteristics of Gaussian distribution, the association weight is calculated according to the relative distance |ji| between the time points. The closer the distance, the greater the weight, which reflects the closer association between adjacent time points in the time series. Rescale[] is the form of converting the association weight matrix into discrete distribution; (2-3) Series correlation is the correlation relationship obtained by adaptive learning from the original time series data. It can capture the dynamic dependency relationship in the time series and reflect the real degree of correlation between time points. The series correlation calculation process is implemented through the standard self-attention mechanism. Its calculation formula is: in represents the output of the l-1th layer, l∈{1, 2, …, L}, is the learnable parameter matrix, the dimensions of Q, K, and V are N*d model , N is the length of the time series, d model is the number of channels of hidden states in the model, QK T Represents the dot product of query Q and key K, and obtains an attention score matrix with dimension N*N; (2-4) In the multi-layer Mixer structure, the prior association PA and series association SA of each layer are calculated; PA and SA are the sets of calculation results of each layer, expressed as: PA={PA 1 ,PA 2 ,…,PA L } IN={IN 1 ,IN 2 ,…,IN L } (2-5) Use the symmetric KL divergence to calculate the difference between the prior association and the series association, while considering and The KL divergence of the i-th time point in the l-th layer is defined as: (2-6) Considering the association information learned at different levels in the multi-layer model, the association differences of each layer are fused to obtain a more representative and stable association difference metric AssDis(PA, SA; x), where L represents the number of layers of the multi-layer model; 5. The multivariate time series adaptive anomaly detection method based on Mixer-Transformer according to claim 4 is characterized in that In the anomaly score calculation module in step (3), the anomaly score of each time point in the time series is calculated by the associated difference; the anomaly score of each time point in the time series is defined as: in It is the reconstruction result of the input time series data χ.
6. The multivariate time series adaptive anomaly detection method based on Mixer-Transformer according to claim 5 is characterized in that The specific steps of step (4) are as follows: (4-1) Introduce the setting of sliding window. At each time t, observe the abnormal score sequence {S t-w , …, S t ,…S t+w }; (5-2) Use the exponentially weighted moving average method to smooth the input values in order to better capture the long-term trends in the data; given the anomaly score S of the current time step t , its corresponding exponentially weighted moving average EWMA t It can be updated through the following recursive relationship: EWMA t =αEWMA t-1 +(1-α)S t EWMA t-1 is the exponentially weighted moving average of the previous moment, α∈(0,1) is the smoothing factor, a larger α value means that the model pays more attention to the most recent data points, while a smaller α value will increase the reliance on historical data, making the model smoother; (5-3) A sensitive response limit threshold is set according to the difference between the current input value and the moving average value. The threshold update formula is as follows: where δ t is the threshold at the current moment, which is used to determine whether the current data deviates significantly from the expected trend. Calculate the current anomaly score S t With the moving average EWMA t The absolute deviation between the current observation and the smoothed trend.
Citation Information
Cited By
Software development application data processing method based on AI large model
CN121433618A