Time series data anomaly detection method and system based on model routing mechanism

Through the timing data anomaly detection method based on the model routing mechanism, hierarchical matching and high-dimensional embedding optimization technology, the problem of degradation of detection accuracy and excessive consumption of computing resources caused by changes in equipment operating conditions is solved, and efficient and robust abnormality detection is achieved.

CN120337078APending Publication Date: 2025-07-18CRRC IND INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510471298.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the existing timing data abnormality detection methods face changes in the operating conditions of the equipment, the strong coupling between the model and the data distribution leads to a decrease in detection accuracy, and the integrated learning method has a large overhead for computing resources, making it difficult to effectively deploy in industrial scenarios with resource limitations.

Method used

The timing data anomaly detection method based on the model routing mechanism is adopted, and the abnormality detection model library is built, and the optimal model unit is selected for detection using the hierarchical matching mechanism, including calculation time-frequency domain statistical feature map, Euclidean distance normalization, uniform distribution loss optimization and high-dimensional embedding to achieve adaptive matching and efficient call of the model.

Benefits of technology

It realizes efficient abnormal detection in different scenarios, reduces the computing resource requirements, maintains detection accuracy, and improves the generalization ability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337078A_ABST
    Figure CN120337078A_ABST
Patent Text Reader

Abstract

The invention provides a time series data anomaly detection method and system based on a model routing mechanism, and the method comprises the steps: S1, independently training an adaptive anomaly detection model unit according to the obtained various historical time series data of equipment in different operation stages and working conditions, and constructing an anomaly detection model library; s2, collecting to-be-detected time series data on line, establishing a model routing mechanism based on hierarchical matching, selecting the anomaly detection model unit, detecting the to-be-detected time series data by the selected anomaly detection model unit, and sending the to-be-detected time series data to a server; and the model routing mechanism quantitatively evaluates the association degree of the time series data to be detected and the historical time series data in two dimensions of time-frequency domain statistical characteristics and characteristic distribution in a layered manner, performs primary selection and further optimization of an anomaly detection model unit, and realizes adaptive matching of an optimal anomaly detection model in a new scene. According to the invention, anomaly detection can be carried out on time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of time series data anomaly detection, and in particular to a time series data anomaly detection method and system based on a model routing mechanism. Background Art

[0002] The rapid development of the Industrial Internet of Things has led to an explosive growth in multi-source sensor time series data such as flow, temperature, and vibration. Anomaly detection based on such data has become the core technology for intelligent operation and maintenance of industrial equipment. The current mainstream time series anomaly detection methods can be summarized into three categories: (1) Threshold-based rule methods, which determine the degree of data deviation by setting static or dynamic thresholds; (2) Unsupervised learning methods, such as mining the potential distribution characteristics of data based on clustering, isolation forest and other algorithms; (3) Self-supervised learning methods, including generative models based on time series reconstruction errors and representation models based on contrastive learning. When facing the stable working conditions of a single device, these methods can achieve high detection accuracy after parameter tuning, but their core defect lies in the strong coupling between the model and data distribution: when the equipment operating conditions switch (such as load fluctuations, environmental changes) or enter different life cycle stages, the statistical characteristics of the time series data (such as mean, variance, spectrum mean, spectrum variance, spectrum peak) will evolve dynamically, resulting in the failure of the preset threshold or static feature extraction logic of a single model, and the need to frequently readjust the model parameters, significantly increasing the operation and maintenance costs.

[0003] In order to improve the generalization and robustness of the model, the ensemble learning method uses the complementarity between models to enhance the dynamic data adaptability by deploying heterogeneous detection algorithms such as threshold rules, clustering models, and deep generative models in parallel. However, due to the lack of sufficient verification of unknown detection scenarios in the offline stage, the performance of each anomaly detection model unit in actual deployment is difficult to reliably evaluate, which makes the integration strategy based on confidence assessment difficult to implement due to the lack of prior knowledge, and the majority voting mechanism may cause the optimal detection results adapted to the current data characteristics to be diluted by the misjudgment of other models. In addition, the high computing resource overhead brought by the parallel execution of multiple models further limits the feasibility of the actual deployment of the integration method in resource-constrained industrial scenarios. Therefore, how to establish a dynamic model selection mechanism for the characteristics of the data to be tested, reduce the redundant computing load while ensuring the detection accuracy, and build a lightweight adaptive integration framework is the core challenge to achieve cross-scenario generalized anomaly detection. Summary of the invention

[0004] In order to solve the technical problems existing in the above-mentioned prior art, the present invention provides a method and system for detecting anomalies in time series data based on a model routing mechanism. The technical solution is as follows:

[0005] On the one hand, a method for detecting anomalies in time series data based on a model routing mechanism is provided, the method comprising:

[0006] S1. Independently train an adapted anomaly detection model unit based on various historical time-series data of the device under different operating stages and working conditions, and construct an anomaly detection model library.

[0007] S2. Online collect the time-series data to be measured, establish a model routing mechanism based on hierarchical matching to select the anomaly detection model unit, and use the selected anomaly detection model unit to detect the time-series data to be measured. The model routing mechanism hierarchically and quantitatively evaluates the correlation degree between the time-series data to be measured and the historical time-series data in two dimensions of time-frequency domain statistical features and feature distribution, conducts preliminary selection and further optimization of the anomaly detection model unit, and realizes the adaptive matching of the optimal anomaly detection model in the new scenario.

[0008] Optionally, the preliminary selection of the anomaly detection model unit in S2 specifically includes:

[0009] S211. Calculate the time-frequency domain statistical features of various historical time-series data, and establish a mapping relationship library from data to models.

[0010] S212. Calculate the time-frequency domain statistical features of the time-series data to be measured.

[0011] S213. Calculate the Euclidean distance between the time-frequency domain statistical features of the time-series data to be measured and the corresponding time-frequency domain statistical features of various historical time-series data, and normalize the obtained Euclidean distance.

[0012] S214. For each type of historical time-series data, perform weighted summation on the normalized Euclidean distances between its time-frequency domain statistical features and the corresponding time-frequency domain statistical features of the time-series data to be measured. The result of the weighted summation represents the difference value between the time-series data to be measured and various historical time-series data in terms of time-frequency domain statistical features. Arrange the top m types of historical time-series data {H1, H2,..., H m} and the corresponding set of anomaly detection model units {M1, M2,..., M m} as candidate historical time-series data and candidate models, where m is an adjustable parameter used to control the scale of the candidate models.

[0013] Optionally, the further optimization of the anomaly detection model unit in S2 specifically includes:

[0014] S221. Perform standardized preprocessing on the time-series data X to be measured and the candidate historical time-series data H i .

[0015] S222. Split and align the features of the preprocessed time-series data to be measured and each type of candidate historical time-series data, respectively obtaining n and s data segments, and each data segment is a high-dimensional vector of c·d dimensions.

[0016] S223. Optimize the high-dimensional time-series data embedding of the data segment based on the uniform distribution loss, and convert the feature dimension from c·d dimensions to k dimensions, so that different candidate historical time-series data are evenly distributed in the new high-dimensional space, and enhance the distribution difference between candidate historical time-series data;

[0017] S224. Randomly select w data segments from the time-series data to be measured and each type of candidate historical time-series data after the feature dimension conversion, and use a variety of distribution difference measurement methods to calculate the multi-index distribution difference;

[0018] S225. After each distribution difference measurement method j, sort each type of candidate historical time-series data in ascending order according to the difference degree under the current method, and generate the serial number R j (H i ) ∈ {1, 2,.., k}, and the serial number of the one with the smallest difference is 1;

[0019] S226. For each type of candidate historical time-series data, obtain the comprehensive ranking score of each type of candidate historical time-series data H i by summing up its serial numbers under each distribution difference measurement method;

[0020]

[0021] S227. Sort all candidate historical time-series data in ascending order according to the comprehensive ranking score, and select the abnormal detection model unit M* corresponding to the candidate historical time-series data H* with the smallest comprehensive ranking score as the matching model unit;

[0022] S228. Load the historical training parameters of the abnormal detection model unit M*, and activate this abnormal detection model unit to be used for detecting the time-series data to be measured.

[0023] Optionally, the S222 specifically includes:

[0024] Perform dimension conversion on the time-series data to be measured and candidate historical time-series data in the way of sliding window segmentation. Set a unified feature dimension d for the time-series data to be measured and candidate historical time-series data, and segment the time-series data in a non-overlapping manner. The time-series data to be measured will generate data segments, N is the length of the time-series data to be measured, the complete length of the historical time-series data is greater than the time-series data to be measured, assuming its data length is S, and it will generate s = S / d data segments, s > n. Before segmentation, the time-series data to be measured X = {x1, x2..., x 10 , x 11 ,..., x 20 ,...}, and after segmentation, X T = {{x1,..., x10 , {x 11 ,..., x 20},...}, denoted as X T = {X T 1 , X T 2 , …}, the candidate historical time - series data H before segmentation i = {h1, h2..., h 10 , h 11 ,..., h 20 ,...}, after segmentation, we get H i,T = {{h1,..., h 10}, {h 11 ,..., h 20},...}, denoted as H i,T = {H i,T 1 , H i,T 2 ,...}, for the time - series data collected by a single - channel sensor, x i and h i are numerical scalars. For the time - series data collected by a multi - channel sensor with c sampling channels, x i and h i are arrays with c numerical values. After data segmentation, the multi - channel values in each window are concatenated into a c·d - dimensional vector.

[0025] Optionally, the S223 specifically includes:

[0026] S2231. Each time, randomly select a data segment from each type of candidate historical data to form a batch of data, obtaining Batch b = {H 1,T b , H 2,T b ..., H m,T b};

[0027] S2232. Construct an encoder f θ : R c·d → R k , mapping the historical time - series data segment to a k - dimensional space, where k > c·d;

[0028] S2233. Perform L2 normalization on the encoder output vector e i = f θ (H i,T j ), and constrain it to the unit hypersphere:

[0029]

[0030] S2234. Uniform distribution loss calculation, including:

[0031] Calculating the Euclidean distance matrix of all data segment embeddings within a batch D i,j = ||e i - e j ||₂, i, j ∈ {1, 2,..., m}, representing the Euclidean distance between two data segment embeddings;

[0032] Calculating the loss function where is the similarity weight generated from the Euclidean distance matrix through Gaussian kernel transformation. Introducing a temperature parameter t to control the distribution smoothness, the smaller the value of the loss function, the more uniform the distribution of every two data segments;

[0033] S2235. Performing feature space transformation, including:

[0034] Using the trained encoder to encode all data segments corresponding to the time series to be measured and the candidate historical time series data, and converting the original feature dimensions of X T and H i,T from c·d dimensions to k dimensions, denoted as X k and H i,k .

[0035] Optionally, the encoder includes, but is not limited to, a deep neural network with one of the following structures:

[0036] Convolutional Neural Network (CNN), extracting local time series features through a one-dimensional convolutional layer, Long Short-Term Memory Network (LSTM), capturing long-term time series dependencies, Transformer encoder, using self-attention mechanism to model global time series patterns;

[0037] The training strategy of the encoder is as follows:

[0038] Training the network parameters of the encoder through the gradient descent optimization algorithm, with the goal of minimizing the loss function. During training, a dynamic temperature adjustment strategy is adopted. In the initial training stage, a lower temperature t = 0.1 is set and gradually increased to t = 1.0 to refine the distribution uniformity. Set the training stop condition. When the loss function value of a batch of input data segments reaches λ, stop training. λ is a hyperparameter and is set according to needs. At the same time, set an early stopping mechanism. If the loss function value does not decrease for p rounds, stop training. p is a hyperparameter and is set according to needs.

[0039] Optionally, the multiple distribution difference measurement methods include, but are not limited to: maximum mean difference, cosine similarity, L2 distance, A-distance, KL divergence, and correlation alignment.

[0040] On the other hand, a time series data anomaly detection system based on a model routing mechanism is provided. The system includes:

[0041] A construction module for independently training and adapting anomaly detection model units according to various historical time series data of the device under different operation stages and working conditions, and constructing an anomaly detection model library;

[0042] A selection module for online collecting the time series data to be measured, establishing a model routing mechanism based on hierarchical matching, selecting the anomaly detection model units, and detecting the time series data to be measured by the selected anomaly detection model units. The model routing mechanism hierarchically and quantitatively evaluates the correlation degree between the time series data to be measured and the historical time series data in two dimensions of time-frequency domain statistical features and feature distribution, conducts primary selection and further optimization of the anomaly detection model units, and realizes the adaptive matching of the optimal anomaly detection model in the new scenario.

[0043] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned time series data anomaly detection method based on the model routing mechanism.

[0044] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned time series data anomaly detection method based on the model routing mechanism.

[0045] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0046] Through the model routing method, the present invention realizes the effective integration and flexible invocation of various anomaly detection algorithms, can adaptively select and match the anomaly detection model units according to the time series data to be measured, realizes the full utilization of historical time series data and the trained anomaly detection models, and is an anomaly detection method with application scenario generalization ability.

[0047] The model routing method of the present invention adopts a hierarchical matching mechanism. First, according to the time-frequency domain statistical features of the time series data, primary selection of historical time series data and corresponding candidate models is carried out, which can significantly reduce the calculation amount of subsequent data distribution differences, and further keep the calculation amount of the entire model routing method controllable.

[0048] Through further optimization, the model routing method in the present invention only selects an anomaly detection model unit with the highest matching degree to perform anomaly detection on the subsequent time series data to be measured. This makes the computational complexity of the anomaly detection method in the present invention at the same level as that of a single anomaly detection method, significantly lower than the traditional anomaly detection method based on ensemble learning. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0050] Figure 1 is a flowchart of a time series data anomaly detection method based on a model routing mechanism provided by an embodiment of the present invention;

[0051] Figure 2 is an overall block diagram of a time series data anomaly detection method based on a model routing mechanism provided by an embodiment of the present invention;

[0052] Figure 3 is a block diagram of a time series data anomaly detection system based on a model routing mechanism provided by an embodiment of the present invention;

[0053] Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0054] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the drawings and specific embodiments.

[0055] An embodiment of the present invention provides a time series data anomaly detection method based on a model routing mechanism. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 As shown in the flowchart of this method, Figure 2 As shown in the overall block diagram of this method, the processing flow can include the following steps:

[0056] S1. According to the various historical time series data of the device under different operating stages and working conditions obtained, independently train the adapted anomaly detection model units, and construct an anomaly detection model library;

[0057] S2. Online collect the time series data to be measured, establish a model routing mechanism based on hierarchical matching, select the anomaly detection model unit, and use the selected anomaly detection model unit to detect the time series data to be measured. The model routing mechanism hierarchically and quantitatively evaluates the correlation degree between the time series data to be measured and the historical time series data in two dimensions: the statistical characteristics and feature distribution in the time-frequency domain, and conducts the primary selection and further optimization of the anomaly detection model unit to achieve the adaptive matching of the optimal anomaly detection model in the new scenario.

[0058] Combined with the actual scenario of the industrial field, the time series data to be measured in the embodiments of the present invention are data collected online at a certain frequency, that is, the time series data to be measured are streaming data (in the embodiments of the present invention, the time series data to be measured for N seconds can be collected online, and the specific collection duration is determined by the sampling period of the sensor and the operation period of the device, and N is set as an adjustable parameter). Correspondingly, the historical time series data are static historical data that have been saved (the time series data to be measured and the historical time series data in the embodiments of the present invention are data corresponding to the same physical quantity, and have the same original dimension and the same sampling frequency. For example, if the time series data to be measured are the vibration state data of the device and have three sampling channels, then the selected historical time series data for comparison should also be vibration state data and have three sampling channels. The length of the historical time series data is longer than the length N of the time series data to be measured. In the primary selection stage of the candidate model, the first N time step data of the historical time series data are intercepted for comparison).

[0059] Optionally, according to the various historical time series data of the device obtained in S1 under different operation stages and working conditions, independently train the adapted anomaly detection model units and construct an anomaly detection model library, specifically including:

[0060] Sort out or collect the historical time series data (such as sensor data of flow, temperature, vibration, etc., which are specifically determined according to the device) of the device under different operation stages (initial operation stage (running-in period), mature operation stage (stable period), decline stage (aging period)) and working conditions (stable operation condition (torque and speed remain constant, and can be subdivided into multiple specific working conditions according to the values of speed and torque), dynamic operation condition (torque or speed is in a fluctuating state, and can be subdivided into multiple specific working conditions according to the fluctuation range (maximum value minus minimum value) of speed and torque)), and construct a historical time series data set;

[0061] Independently train the adapted anomaly detection model units for each type of (different operation stages + different working conditions) historical data to form an anomaly detection model library composed of heterogeneous algorithm units including rule-based methods based on thresholds, unsupervised learning methods, and self-supervised learning methods, etc.

[0062] In an embodiment of the present invention, the anomaly detection model unit refers to an anomaly detection algorithm adapted to a certain historical time-series data, which includes but is not limited to a specific algorithm among rule-based methods based on thresholds, unsupervised learning methods, or self-supervised learning methods, and has completed a training process based on this historical time-series data and has the function of performing anomaly detection on this historical time-series data.

[0063] Optionally, the primary selection of the anomaly detection model unit in S2 specifically includes:

[0064] S211. Calculate various time-frequency domain statistical features (such as time-domain mean, time-domain variance, frequency-spectrum mean, frequency-spectrum variance, frequency-spectrum peak, etc.) of various historical time-series data, and establish a mapping relationship library from data to the model;

[0065] S212. Calculate various time-frequency domain statistical features (such as time-domain mean, time-domain variance, frequency-spectrum mean, frequency-spectrum variance, frequency-spectrum peak, etc.) of the to-be-tested time-series data;

[0066] S213. Calculate the Euclidean distance between the to-be-tested time-series data and various corresponding time-frequency domain statistical features of various historical time-series data. Taking the time-domain mean of the to-be-tested time-series data X and the historical time-series data M i as an example, they are μ X and μ Y,i , respectively, then their Euclidean distance is and normalize the obtained Euclidean distance (to eliminate the differences in the dimensions of different statistical features);

[0067] S214. For each type of historical time-series data, perform weighted summation on the normalized Euclidean distances between its various time-frequency domain statistical features and the corresponding time-frequency domain statistical features of the to-be-tested time-series data (each weight can be set as needed). The result of the weighted summation represents the difference value between the to-be-tested time-series data and various historical time-series data in terms of time-frequency domain statistical features. Arrange the top m types of historical time-series data {H1, H2,..., H m} and the corresponding anomaly detection model unit set {M1, M2,..., M m} as candidate historical time-series data and candidate models, where m is an adjustable parameter used to control the scale of the candidate models.

[0068] Optionally, the further optimization of the anomaly detection model unit in S2 specifically includes:

[0069] S221. Perform standardized preprocessing (which can be a normalization operation) on the to-be-tested time-series data X and the candidate historical time-series data H i ;

[0070] S222. Split and align the features of the preprocessed time series data to be measured and each type of candidate historical time series data, respectively obtaining n and s data segments, each data segment being a high-dimensional vector of c·d dimensions;

[0071] Optionally, the S222 specifically includes:

[0072] Perform dimensionality conversion on the time series data to be measured and the candidate historical time series data by means of sliding window segmentation. Set a unified feature dimension d (for example, d = 10) for the time series data to be measured and the candidate historical time series data, and segment the time series data in a non-overlapping manner (that is, the moving step of the sliding window is equal to the size of the sliding window, which is the set feature dimension d). The time series data to be measured will generate data segments, where N is the length of the time series data to be measured. The complete length of the historical time series data is greater than that of the time series data to be measured. Assume its data length is S, and it will generate s = S / d data segments, s > n. Before segmentation, the time series data to be measured X = {x1, x2..., x 10 , x 11 ,..., x 20 ,...}, and after segmentation, X T ={{x1,..., x 10}, {x 11 ,..., x 20},...}, denoted as X T ={X T 1 , X T 2 ,...}. Before segmentation, the candidate historical time series data H i ={h1, h2..., h 10 , h 11 ,..., h 20 ,...}, and after segmentation, H i,T ={{h1,..., h 10}, {h 11 ,..., h 20},...}, denoted as H i,T ={H i,T 1 , H i,T 2 ,...}. For the time series data collected by a single-channel sensor, x i and h i are numerical scalars. For the time series data collected by a multi-channel sensor with c sampling channels, x i and h i are arrays with c numerical values. After data segmentation, the multi-channel values in each window are concatenated into a c·d-dimensional vector.

[0073] S223. Optimize the high-dimensional time-series data embedding of the data segment based on the uniform distribution loss, and convert the feature dimension from c·d dimensions to k dimensions, so that different candidate historical time-series data are evenly distributed in the new high-dimensional space, and the difference in distribution among candidate historical time-series data is enhanced;

[0074] The matching between the time-series data to be measured and multiple candidate historical time-series data may result in a decrease in the matching accuracy due to the overlapping distribution of historical time-series data. Therefore, the embodiments of the present invention propose an unsupervised high-dimensional embedding optimization method. By designing the loss function of the encoder in the training stage, the distribution distance of different candidate historical time-series data in the high-dimensional space is enlarged (more evenly distributed), thereby significantly enhancing the difference in distribution among historical time-series data and reducing the problem of the decrease in the matching accuracy of the time-series data to be measured caused by the overlapping distribution of historical time-series data.

[0075] Optionally, S223 specifically includes:

[0076] S2231. Randomly select a data segment from each type of candidate historical data each time to form a batch of data, and obtain Batch b ={H 1,T b ,H 2,T b ...,H m,T b};

[0077] S2232. Construct an encoder f θ :R c·d →R k , and map the historical time-series data segment to the k-dimensional space, where k>c·d;

[0078] S2233. Perform L2 normalization on the encoder output vector e i =f θ (H i,T j ), and constrain it to the unit hypersphere:

[0079]

[0080] S2234. Calculate the uniform distribution loss, including:

[0081] Calculate the Euclidean distance matrix of the embeddings of all data segments within the batch D i,j =||e i -e j ||2, i,j∈{1,2,...,m}, representing the Euclidean distance between the embeddings of two data segments;

[0082] Calculate the loss function Among them is the similarity weight generated from the Euclidean distance matrix through Gaussian kernel transformation. A temperature parameter t is introduced to control the distribution smoothness. The more evenly distributed every two data segments are, the smaller the value of the loss function;

[0083] S2235. Perform feature space transformation, including:

[0084] Use the trained encoder to encode all data segments corresponding to the time series to be measured and the candidate historical time series data, and convert the original feature dimensions of X T and H i,T from c·d dimensions to k dimensions, denoted as X k and H i,k respectively.

[0085] Optionally, the encoder includes, but is not limited to, a deep neural network with one of the following structures:

[0086] Convolutional Neural Network (CNN), which extracts local time series features through a one-dimensional convolutional layer, Long Short-Term Memory Network (LSTM), which captures long-term time series dependencies, and Transformer encoder, which models global time series patterns using self-attention mechanism;

[0087] The training strategy of the encoder is as follows:

[0088] Train the network parameters of the encoder through the gradient descent optimization algorithm. The goal is to minimize the loss function. During training, a dynamic temperature adjustment strategy is adopted. In the initial training stage, a lower temperature t = 0.1 is set and gradually increased to t = 1.0 to refine the distribution uniformity. Set the training stop condition. When the loss function value of a batch of input data segments reaches λ, stop training. λ is a hyperparameter and is set according to needs. At the same time, set an early stopping mechanism. If the loss function value does not decrease for p rounds, stop training. p is a hyperparameter and is set according to needs.

[0089] S224. Randomly select w data segments from the time series to be measured after feature dimension transformation and each type of candidate historical time series data, and use multiple distribution difference measurement methods to calculate the multi-index distribution difference;

[0090] Optionally, the multiple distribution difference measurement methods include, but are not limited to: Maximum Mean Discrepancy, Cosine Similarity, L2 Distance, A-distance, KL Divergence, and Correlation Alignment.

[0091] S225. After each distribution difference measurement method j is performed, sort each type of candidate historical time series data in ascending order according to the difference degree under the current method, and generate the serial number R j (H i) ∈ {1, 2, .., k}, the serial number of the one with the smallest difference is 1;

[0092] For example, if in the calculation of the maximum mean difference (MMD), the distribution difference between H 3,k and X k is the smallest, then R MMD (H 3,k ) = 1.

[0093] S226. For each type of candidate historical time series data, by summing up its serial numbers under various distribution difference measurement methods, obtain the comprehensive ranking score of each type of candidate historical time series data H i :

[0094]

[0095] S227. Arrange all candidate historical time series data in ascending order of the comprehensive ranking score, and select the abnormal detection model unit M* corresponding to the candidate historical time series data H* with the smallest comprehensive ranking score as the matching model unit;

[0096] S228. Load the historical training parameters (such as threshold intervals, clustering centers, network weights, etc.) of the abnormal detection model unit M*, and activate this abnormal detection model unit to detect the time series data to be measured.

[0097] For the abnormal detection scenario of industrial field on-line streaming time series data, detecting the time series data to be measured includes the following two sub-steps:

[0098] Data buffer pool management: Pre-define the length L (number of data sampling points) of the time series data buffer pool. When the real-time data stream accumulates to L, update the buffer in accordance with the first-in-first-out mechanism;

[0099] Detection task scheduling: Configure the execution period T of the abnormal detection algorithm, and perform detection on the data in the time series data buffer pool through a periodic trigger mechanism, and output the detection conclusion in real time.

[0100] As Figure 3 shown, the embodiment of the present invention also provides a time series data abnormal detection system based on a model routing mechanism, and the system includes:

[0101] A construction module 310, configured to independently train and adapt abnormal detection model units according to various historical time series data of the device under different operation stages and working conditions, and construct an abnormal detection model library;

[0102] A selection module 320 is configured to collect online the time-series data to be measured, establish a model routing mechanism based on hierarchical matching to select the anomaly detection model units, and use the selected anomaly detection model units to detect the time-series data to be measured. The model routing mechanism hierarchically and quantitatively evaluates the correlation degree between the time-series data to be measured and the historical time-series data in two dimensions of time-frequency domain statistical features and feature distribution, performs primary selection and further optimization of the anomaly detection model units, and realizes the adaptive matching of the optimal anomaly detection model in the new scenario.

[0103] The time-series data anomaly detection system based on a model routing mechanism provided by an embodiment of the present invention has a functional structure corresponding to the time-series data anomaly detection method based on a model routing mechanism provided by an embodiment of the present invention, and will not be elaborated herein.

[0104] Figure 4 FIG. is a schematic structural diagram of an electronic device 400 provided by an embodiment of the present invention. The electronic device 400 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 401 and one or more memories 402. Among them, at least one instruction is stored in the memory 402, and the at least one instruction is loaded and executed by the processor 401 to implement the steps of the above-mentioned time-series data anomaly detection method based on a model routing mechanism.

[0105] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above-mentioned time-series data anomaly detection method based on a model routing mechanism. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0106] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.

[0107] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting anomalies in time series data based on a model routing mechanism, characterized in that The method includes: S1. Independently train an adapted anomaly detection model unit based on various historical time-series data of the device under different operating stages and working conditions, and construct an anomaly detection model library. S2. Online collect the time-series data to be measured, establish a model routing mechanism based on hierarchical matching, select the anomaly detection model unit, and use the selected anomaly detection model unit to detect the time-series data to be measured. The model routing mechanism hierarchically and quantitatively evaluates the correlation degree between the time-series data to be measured and the historical time-series data in two dimensions of time-frequency domain statistical features and feature distribution, conducts the primary selection and further optimization of the anomaly detection model unit, and realizes the adaptive matching of the optimal anomaly detection model in the new scenario.

2. The method according to claim 1, wherein In S2, the primary selection of the anomaly detection model unit specifically includes: S211. Calculate various time-frequency domain statistical features of various historical time-series data, and establish a mapping relationship library from data to models. S212. Calculate various time-frequency domain statistical features of the time-series data to be measured. S213. Calculate the Euclidean distance between the time-frequency domain statistical features corresponding to the time-series data to be measured and various historical time-series data, and normalize the obtained Euclidean distance. S214. For each type of historical time-series data, perform a weighted summation of the normalized Euclidean distances between its time-frequency domain statistical features and the corresponding time-frequency domain statistical features of the time-series data to be measured. The result of the weighted summation represents the difference value between the time-series data to be measured and each type of historical time-series data in terms of time-frequency domain statistical features. Arrange the first m types of historical time-series data {H1, H2,..., H m} and the corresponding set of anomaly detection model units {M1, M2,..., M m} as the candidate historical time-series data and candidate models, where m is an adjustable parameter used to control the scale of the candidate models.

3. The method according to claim 2, wherein In S2, the further optimization of the anomaly detection model unit specifically includes: S221. Perform standard preprocessing on the to-be-tested time series data X and the candidate historical time series data H i ; S222. Segment and align the preprocessed time-series data to be measured and each type of candidate historical time-series data to obtain n and s data segments respectively, and each data segment is a high-dimensional vector of c·d dimensions. S223. Optimize the embedding of high-dimensional time-series data based on the uniform distribution loss for the data segments, convert the feature dimension from c·d dimensions to k dimensions, so that different candidate historical time-series data are evenly distributed in the new high-dimensional space, and improve the distribution difference between candidate historical time-series data. S224. Randomly select w data segments from the time-series data to be measured and each type of candidate historical time-series data after the feature dimension conversion, and use a variety of distribution difference measurement methods to calculate the multi-index distribution difference. S225. After each distribution difference measurement method j is performed, each category of candidate historical time series data is sorted in ascending order according to the degree of difference under the current method to generate a serial number R j (H i ) ∈ {1, 2,.., k}, and the serial number of the one with the smallest difference is 1; S226. For each type of candidate historical time-series data, by summing up its serial numbers under various distribution difference measurement methods, the comprehensive ranking score of each type of candidate historical time-series data H i is obtained: S227. Arrange all candidate historical time-series data in ascending order of the comprehensive sorting score, and select the anomaly detection model unit M* corresponding to the candidate historical time-series data H* with the smallest comprehensive sorting score as the matching model unit. S228. Load the historical training parameters of the anomaly detection model unit M*, and activate this anomaly detection model unit to detect the time-series data to be measured.

4. The method according to claim 3, wherein S222 specifically includes: The sliding window segmentation method is used to perform dimensionality conversion on the time series data to be measured and the candidate historical time series data. A unified feature dimension d is set for the time series data to be measured and the candidate historical time series data. The time series data is segmented in a non-overlapping manner. The time series data to be measured will generate data segments. N is the length of the time series data to be measured. The complete length of the historical time series data is greater than that of the time series data to be measured. Assume its data length is S, and it will generate s = S / d data segments, where s > n. Before segmentation, the time series data to be measured X = {x1, x2..., x 10 , x 11 ,..., x 20 ,...}. After segmentation, we get X T = {{x1,..., x 10}, {x 11 ,..., x 20},...}, denoted as X T = {X T 1 , X T 2 ,...}. Before segmentation, the candidate historical time series data H i = {h1, h2..., h 10 , h 11 ,..., h 20 ,...}. After segmentation, we get H i,T = {{h1,..., h 10}, {h 11 ,..., h 20},...}, denoted as H i,T = {H i,T 1 , H i,T 2 ,...}. For the time series data collected by a single-channel sensor, x i and h i are numerical scalars. For the time series data collected by a multi-channel sensor with c sampling channels, x i and h i are arrays with c numerical values. After data segmentation, the multi-channel values in each window are concatenated into a c·d-dimensional vector.

5. The method according to claim 4, wherein S223 specifically includes: S2231. Randomly select a data segment from each type of candidate historical data each time to form a batch of data, obtaining Batch b ={H 1,T b , H 2,T b ..., H m,T b}; S2232. Construct the encoder f θ : R c·d →R k , map the historical time series data segment to a k-dimensional space, where k > c·d; S2233. Perform L2 normalization on the encoder output vector e i = f θ (H i,T j ) and constrain it to the unit hypersphere: S2234. Uniform distribution loss calculation, including: Calculate the Euclidean distance matrix of all data segment embeddings within a calculation batch D i,j = ||e i - e j ||2, i, j ∈ {1, 2,..., m}, representing the Euclidean distance between two data segment embeddings; Calculate the loss function where is the similarity weight generated from the Euclidean distance matrix through Gaussian kernel transformation. A temperature parameter t is introduced to control the distribution smoothness. The more evenly distributed every two data segments are, the smaller the value of the loss function is; S2235. Conduct feature space conversion, including: Using the trained encoder, encode all data segments corresponding to the time series data to be measured and the candidate historical time series data, and convert the original feature dimensions of X T and H i,T from c·d dimensions to k dimensions, denoted as X k and H i,k .

6. The method according to claim 5, characterized in that The encoder includes, but is not limited to, a deep neural network with one of the following structures: Convolutional neural network CNN, which extracts local time-series features through a one-dimensional convolutional layer, long short-term memory network LSTM, which captures long-term time-series dependencies, and Transformer encoder, which models global time-series patterns using the self-attention mechanism. The training strategy of the encoder is as follows: The network parameters of the encoder are trained by the gradient descent optimization algorithm, with the goal of minimizing the loss function. During training, a dynamic temperature adjustment strategy is adopted. In the initial training stage, a lower temperature t = 0.1 is set and gradually increased to t = 1.0 to refine the distribution uniformity. The stopping condition of training is set. When the value of the loss function for a certain batch of input data segments reaches λ, training stops. λ is a hyperparameter and is set as needed. At the same time, an early stopping mechanism is set. If the value of the loss function does not decrease for p rounds, training stops. p is a hyperparameter and is set as needed.

7. The method according to claim 3, characterized in that The multiple distribution difference measurement methods include, but are not limited to: maximum mean difference, cosine similarity, L2 distance, A-distance, KL divergence, and correlation alignment.

8. A time-series data anomaly detection system based on a model routing mechanism, characterized in that, The system includes: A construction module for independently training an adapted anomaly detection model unit according to various historical time-series data of the device under different operating stages and working conditions, and constructing an anomaly detection model library; A selection module for online collecting the time-series data to be measured, establishing a model routing mechanism based on hierarchical matching, selecting the anomaly detection model unit, and detecting the time-series data to be measured by the selected anomaly detection model unit. The model routing mechanism hierarchically and quantitatively evaluates the correlation degree between the time-series data to be measured and the historical time-series data in two dimensions of time-frequency domain statistical features and feature distribution, conducts the initial selection and further optimization of the anomaly detection model unit, and realizes the adaptive matching of the optimal anomaly detection model in the new scenario.

9. An electronic device, the electronic device includes a processor and a memory, and at least one instruction is stored in the memory, characterized in that, The at least one instruction is loaded and executed by the processor to implement the time-series data anomaly detection method based on the model routing mechanism according to any one of claims 1-7.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that, The at least one instruction is loaded and executed by the processor to implement the time-series data anomaly detection method based on the model routing mechanism according to any one of claims 1-7.