Cloud system anomaly detection method based on time sequence classification automatic matching

Through edge computing probe construction and classification of time series, combined with the matching of the abnormality detection algorithm of the master node, the problem of accurate classification of composite type time series data in the cloud system is solved, and data processing efficiency and automation are improved.

CN120342833APending Publication Date: 2025-07-18STATE GRID INFORMATION & TELECOMM BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510656846.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In cloud systems, with the explosive growth of data volume, it is difficult for the existing technology to accurately classify compound type time series data, resulting in low efficiency in intelligent operation and maintenance data processing and insufficient automation.

Method used

The first-level time series is constructed through edge computing probes, feature extraction and preliminary classification are performed, and the second-level time series is constructed for different coding and data compression are performed. Anomaly detection algorithms with different parameters are configured on the main node, and algorithm matching is performed based on the head metadata of the second-level time series to achieve refined detection.

Benefits of technology

It improves the real-time data stream processing capability and automation of the cloud system, reduces the data processing volume of the master node, and improves the accuracy and efficiency of abnormal detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342833A_ABST
    Figure CN120342833A_ABST
Patent Text Reader

Abstract

The invention discloses an on-cloud system anomaly detection method for time sequence classification automatic matching, and relates to a data processing technology, and the method comprises the steps: collecting corresponding node parameters through employing an edge calculation probe, so as to construct a first-stage time sequence; at each secondary node, performing feature extraction and preliminary classification on the primary time sequence to construct a secondary time sequence based on the extracted features and a preliminary classification result; differential coding and data compression are carried out on the constructed secondary time sequence, and then the secondary time sequence is sent to the main node; at the main node, abnormal detection algorithms of different parameters are configured in advance, and algorithm matching is carried out based on header metadata of the secondary time sequence obtained through recognition; and detecting the secondary time sequence according to a matched anomaly detection algorithm to obtain a detection result. According to the invention, the processing capability of the real-time data stream of the cloud system can be improved, and the automation degree of the cloud system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular to an abnormal detection method for a cloud system with automatic matching of time series classification. Background Art

[0002] Time series classification refers to classifying time series data in the operation and maintenance field according to its own statistical characteristics or different components (trend components, periodic components, noise, etc.). The significance of time series classification can be summarized into the following five points: 1) Matching the classified data with a prediction algorithm to improve prediction accuracy and reduce the computational amount when finding the optimal model at the same time; 2) Matching the classified data with an anomaly detection algorithm to improve anomaly detection accuracy and also reduce the computational amount when finding the optimal model; 3) It can be used as a preprocessing step for time series prediction algorithms, such as identifying and removing periodicity in data; 4) It can be used as a preprocessing step for time series anomaly detection algorithms, such as periodic removal in bank batch processing tasks, etc.; 5) Statistical analysis of time series data, such as counting the respective proportions of different types of data in a database, analyzing the importance of different business data, and data quality assessment, etc.

[0003] In time series classification, the real operation and maintenance data may be of a composite type, rather than a simple periodic type, but a periodic type with a trend. In this case, incorrect results will occur if periodic determination is carried out first. With the explosive growth of data volume, intelligent operation and maintenance data processing and analysis technologies are facing challenges in storing, processing, and analyzing massive data. Therefore, accurate classification of time series is required. Summary of the Invention

[0004] An embodiment of this application provides an abnormal detection method for a cloud system with automatic matching of time series classification, aiming to propose an intelligent operation and maintenance data processing and classification method for time series, improve the processing ability of real-time data streams, and improve the automation level of cloud systems.

[0005] An embodiment of this application proposes an abnormal detection method for a cloud system with automatic matching of time series classification. The cloud system includes at least one master node and multiple slave nodes, and the master node and the slave nodes are communicatively connected. Edge computing probes are provided at the slave nodes. The abnormal detection method includes: Collecting corresponding node parameters by using the edge computing probes to construct a first-level time series, where the first-level time series has a preset fixed format; At each secondary node, feature extraction and preliminary classification are performed on the first-level time series, and a second-level time series is constructed based on the extracted features and preliminary classification results, where the second-level time series is obtained by performing position transformation on the preliminary classification results based on the fixed format, and the greater the probability proportion of any parameter being classified into the corresponding category and the more forward the position; The constructed second-level time series is sent to the master node after differential encoding and data compression; At the master node, anomaly detection algorithms with different parameters are preconfigured, and algorithm matching is performed based on identifying the header metadata of the obtained second-level time series; The second-level time series is detected according to the matched anomaly detection algorithm to obtain a detection result.

[0006] Optionally, using edge computing probes to collect corresponding node parameters to construct the first-level time series includes: when a certain parameter is not included in the node parameters collected at any moment, filling with a specified character, and; Constructing the second-level time series also includes: Setting a representative header sequence for each classification, combining the sequence obtained by performing position transformation on the preliminary classification results based on the fixed format with each representative header sequence, and filling the missing parameters with a specified character.

[0007] Optionally, the features extracted from the first-level time series at each secondary node include: Time-domain features, including mean, variance, kurtosis, and density of mutation points; Frequency-domain features, which are the main frequency components and energy distribution extracted by fast Fourier transform; Morphological features, which are image features extracted based on the image expression of parameters; And dynamically adjusting the weights of different features according to the corresponding business scenario.

[0008] Optionally, the preliminary classification of the first-level time series at each secondary node includes: For the first-level time series, data is slid using a sliding window; Calculate the Euclidean distance between all data points within the window; Based on the calculated Euclidean distance, determine candidate values of the neighborhood radius eps, and select the optimal eps that stabilizes the density within the cluster to adjust the eps of the sliding window; Semantic labels are assigned to the features slid by the adjusted sliding window to complete the preliminary classification.

[0009] Optionally, sending the constructed second-level time series to the master node after differential encoding and data compression includes: Map each semantic tag to a binary code as the classification code; Embed the classification code and additional key statistics into the metadata at the head of the secondary time series obtained by performing position transformation; Only use the classification code, additional key statistics, and representative data points of the head sequence as transmission data, and send the transmission data to the master node after compression processing.

[0010] Optionally, the anomaly detection algorithms with different pre-configured parameters include: For periodic data, configure the detection algorithm of seasonal trend decomposition STL + kσ rule, and the configured parameters include the period length and confidence interval threshold, where σ represents the standard deviation, which is used to characterize the dispersion degree of the time series within the window; For mutant data, adopt the cumulative sum control chart CUSUM algorithm, and the configured parameters include the baseline mean and offset threshold; For low signal-to-noise ratio data, use wavelet transform to suppress the data and detect it through the isolation forest, and the configured parameters include the number of wavelet decomposition layers and the number of trees; For trend data, adopt the LSTM prediction model, and the configured parameters include the sliding window size and prediction step length.

[0011] Optionally, based on the recognition of the metadata at the head of the obtained secondary time series, the algorithm matching includes: Establish the matching relationship between the label and the algorithm according to the classification code embedded in the head of the secondary time series.

[0012] Optionally, it further includes: The master node feeds back the anomaly detection result to the secondary node to record the feedback data at each secondary node and optimize the semantic label result.

[0013] The embodiment of the present application also proposes an anomaly detection system for a cloud system with automatic matching of time series classification, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the anomaly detection method for the cloud system with automatic matching of time series classification as described above are implemented.

[0014] The embodiment of the present application proposes an intelligent operation and maintenance data processing classification method for a time series of a cloud system, which improves the processing throughput of real-time data streams and the automation degree of the cloud system.

[0015] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Brief Description of the Drawings

[0016] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a schematic diagram of the basic process of the cloud system anomaly detection method of this embodiment. Detailed Embodiments

[0017] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0018] An embodiment of the present application proposes a cloud system anomaly detection method for automatic matching of time series classification. The cloud system includes at least one master node and multiple slave nodes, and the master node and the slave nodes are communicatively connected. An edge computing probe is provided at the slave nodes. In some specific examples, nodes and links are identified, and the collected data is analyzed and processed to identify each node in the system (such as servers, containers, microservices, etc.) and the links between the nodes (i.e., communication connections). By analyzing information such as the communication traffic and call relationships between nodes, it is determined whether there is a connection between nodes and the strength and stability of the connection.

[0019] Topology generation, based on the identified node and link information, constructs the topology of the system. A graph structure in graph theory can be used to represent the topology, and the master node is specified according to the actual operating conditions. Since the cloud information system has the characteristic of dynamic change, a dynamic update mechanism needs to be established to scan and collect data and monitor the anomalies of nodes and links in the system. As Figure 1 shown, the anomaly detection method of the embodiment of the present application includes: In step S101, an edge computing probe is used to collect corresponding node parameters to construct a first-level time series, where the first-level time series has a preset fixed format. Specifically, metrics such as the response time, throughput (TPS / QPS), error rate, CPU occupancy, CPU process status, memory usage, and network latency of system nodes can be collected and a first-level time series can be constructed. In the embodiment of the present application, after collection, according to a fixed format, that is, each constructed first-level time series has the same metric format.

[0020] In step S102, at each sub-node, feature extraction and preliminary classification are performed on the first-level time series, so as to construct a second-level time series based on the extracted features and the preliminary classification results, where the second-level time series is obtained by performing position transformation on the preliminary classification results based on the fixed format, and the greater the proportion of any parameter in the probability of being classified into the corresponding category and the more forward the position. In a specific example, through preliminary classification and position transformation of the preliminary classification results, the sequence head of the obtained second-level time series can effectively describe the key data of the second-level time series, while other normal data does not appear at the head of the second-level time series.

[0021] In step S103, the constructed second-level time series is subjected to differential encoding and data compression and then sent to the master node. In a specific example, the data volume can be greatly compressed by using the data processing method designed in this application, reducing the data processing volume of the master node.

[0022] In step S104, at the master node, anomaly detection algorithms with different parameters are pre-configured, and algorithm matching is performed based on identifying the header metadata of the obtained second-level time series. By configuring anomaly detection algorithms with different parameters at the master node, the key data of the second-level time series can be refined and detected specifically. In this way, the master node does not need to process the huge data volume of the first-level time series, which can reduce the data processing volume of the master node and improve the efficiency.

[0023] In step S105, the second-level time series is detected according to the matched anomaly detection algorithm to obtain a detection result.

[0024] The embodiment of this application proposes an intelligent operation and maintenance data processing classification method for time series in a cloud system, which improves the ability of real-time data flow and the automation degree of the cloud system.

[0025] In some embodiments, using an edge computing probe to collect corresponding node parameters to construct a first-level time series includes: when a certain parameter is not included in the node parameters collected at any moment, filling with a specified character, for example, in the way of specifying the character *.

[0026] Constructing the second-level time series further includes: setting a representative header sequence for each classification, combining the sequence obtained by performing position transformation on the preliminary classification results based on the fixed format with each representative header sequence, and filling the missing parameters with a specified character. In some examples, a sequence may be classified into multiple categories, then based on the representative header sequence, for the sequence obtained after position transformation, for the missing parameters, fill with a specified character. By this way, the constructed second-level time series has a fixed parameter order under any classification, which is convenient for subsequent feature extraction and model recognition.

[0027] In some embodiments, at each secondary node, the features extracted from the first-level time series include: Time-domain features, including mean, variance, kurtosis, and density of mutation points.

[0028] Frequency-domain features, which are the main frequency components and energy distribution extracted by fast Fourier transform.

[0029] Morphological features, which are image features extracted based on parameter-based image expressions, such as trend slope, periodic intensity, etc.

[0030] In addition, the weights of different features are dynamically adjusted according to the corresponding business scenarios. In a specific example, the business scenarios may include high-load periods, low-load periods, etc. to dynamically adjust the feature weights in the first-level time series.

[0031] In some embodiments, at each secondary node, the preliminary classification of the first-level time series further includes automatically adjusting the neighborhood radius eps of the window based on the distribution density of the data within the window, specifically including: For the first-level time series, a sliding window is used to slide and extract data.

[0032] Calculate the Euclidean distance between all data points within the window; Based on the calculated Euclidean distance, candidate values of the neighborhood radius eps are determined, and the optimal eps that stabilizes the density within the cluster is selected to adjust the eps of the sliding window. For example, the k-distance graph can be used to determine the candidate values of eps, and the optimal eps is selected. In a specific example, the eps value of the historical window can also be introduced for weighted smoothing to avoid parameter mutations caused by instantaneous fluctuations.

[0033] Semantic labels are assigned to the features extracted by the adjusted sliding window to complete the preliminary classification. In a specific example, the semantic labels can be, for example, the preliminary classification of the data within the window, such as periodic type, mutation type, trend type, etc. as semantic labels.

[0034] In some embodiments, after differential encoding and data compression of the constructed second-level time series, sending it to the master node includes: Mapping each semantic label to a binary code as the classification code. For example, configure 0001 to represent the periodic type, 0010 to represent the mutation type, 0011 to represent the trend type, etc., and perform binary mapping, thereby reducing the amount of data sent.

[0035] Embed the classification code into the secondary time series header metadata obtained by performing position transformation, and additionally attach key statistics. In a specific example, for instance, the initially classified periodic type - 0001 is embedded into the secondary time series header metadata, and key statistics such as the cluster mean and extreme values of a sliding window are attached to the header metadata.

[0036] Only use the classification code, the additionally attached key statistics, and the representative data points of the header sequence as transmission data, and after compressing the transmission data, send it to the master node. In a specific example, the representative data points of the header sequence can be obtained by sampling the front data segment, and the representative data points are used to describe the data points that may be abnormal in the current secondary time series.

[0037] In some embodiments, the anomaly detection algorithms with different pre - configured parameters include: For periodic data, configure the detection algorithm of seasonal trend decomposition STL + kσ rule. The configured parameters include the period length and the confidence interval threshold, where σ represents the standard deviation, which is used to characterize the dispersion degree of the time series within the window. The confidence interval threshold is used to define the boundary range for anomaly determination, which is the mean ± k×σ, where k is a multiple. For example, k can be taken as 3, that is, 3σ, and k = 3 can cover 99.7% of the normally distributed data.

[0038] For mutant data, adopt the cumulative sum control chart CUSUM algorithm, and the configured parameters include the baseline mean and the offset threshold.

[0039] For low - signal - to - noise data, use wavelet transform to suppress the data and detect it through an isolation forest. The configured parameters include the number of wavelet decomposition layers and the number of trees.

[0040] For trend - type data, adopt the LSTM prediction model, and the configured parameters include the sliding window size and the prediction step length, where the sliding window size is the window size adaptively adjusted according to the foregoing.

[0041] In some embodiments, based on the recognition of the header metadata of the obtained secondary time series, algorithm matching includes: establishing the matching relationship between the label and the algorithm according to the classification code embedded in the secondary time series header. In some examples, for instance, 0001 (periodic type) → STL decomposition algorithm, 0100 (mutant type) → CUSUM algorithm; 1000 (low signal - to - noise) → wavelet denoising + isolation forest, etc. Further extract the features of the additionally received key statistics and the representative data points of the header sequence, and use the adapted detection method for anomaly detection, thereby greatly improving the accuracy of anomaly recognition.

[0042] In some embodiments, it further includes: the master node feeds back the anomaly detection result to the slave nodes, so as to record the feedback data at each slave node and optimize the semantic label result.

[0043] The solution of this application preliminarily classifies the primary time series at each slave node, and performs algorithm matching and recognition at the master node through encoding. After the preliminary classification at the secondary node, the time series is adjusted so that the data at the head end of the secondary time series contains key data for anomaly detection, and the data is compressed and sent to the master node for accurate detection. The method of this application effectively improves the processing throughput of real-time data streams and improves the automation level of the cloud system.

[0044] An embodiment of this application also proposes an anomaly detection system for a cloud system with automatic matching of time series classification, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the anomaly detection method for the cloud system with automatic matching of time series classification as described above are implemented.

[0045] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present disclosure having equivalent elements, modifications, omissions, combinations (e.g., solutions that cross various embodiments), adaptations, or changes. It is not limited to the examples described in this specification or during the implementation of this application, and the examples will be interpreted as non-exclusive.

[0046] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description.

[0047] The above embodiments are only exemplary embodiments of the present disclosure. Those skilled in the art can make various modifications or equivalent replacements to the present invention within the essence and protection scope of the present disclosure, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.

Claims

1. An abnormal detection method for cloud systems with automatic matching of time series classification, characterized in that The cloud system described above includes at least one master node and multiple slave nodes. The master node and the slave nodes are communicatively connected. An edge computing probe is provided at the slave nodes. The anomaly detection method includes: Using the edge computing probe to collect corresponding node parameters to construct a first-level time series, where the first-level time series has a preset fixed format; At each slave node, perform feature extraction and preliminary classification on the first-level time series to construct a second-level time series based on the extracted features and the preliminary classification results. The second-level time series is obtained by performing position transformation on the preliminary classification results based on the fixed format, and the greater the probability that any parameter is classified into the corresponding category and the more forward the position; Perform differential encoding and data compression on the constructed second-level time series and then send it to the master node; At the master node, pre-configure anomaly detection algorithms with different parameters, and perform algorithm matching based on identifying the header metadata of the obtained second-level time series; According to the matched anomaly detection algorithm, detect the second-level time series to obtain a detection result.

2. The method for anomaly detection of a cloud system with automatic matching of time series classification according to claim 1, wherein Using the edge computing probe to collect corresponding node parameters to construct a first-level time series includes: when the node parameters collected at any moment do not include a certain parameter, filling it with a specified character, and; Constructing the second-level time series further includes: Setting a representative header sequence for each classification, combining the sequence obtained by performing position transformation on the preliminary classification results based on the fixed format with each representative header sequence, and filling the missing parameters with specified characters.

3. The method for detecting cloud system anomalies by automatic matching of time series classification as described in claim 1, characterized in that, At each slave node, the features extracted from the first-level time series include: Time domain features, including mean, variance, kurtosis, and mutation point density; Frequency domain features, which are the main frequency components and energy distribution extracted by fast Fourier transform; Morphological features, which are image features extracted based on the image expression of parameters; And dynamically adjust the weights of different features according to the corresponding business scenarios.

4. The method for detecting cloud system anomalies by automatic matching of time series classification according to claim 3, wherein, At each slave node, the preliminary classification of the first-level time series includes: For the first-level time series, use a sliding window to slide and extract data; Calculate the Euclidean distance between all data points within the window; Based on the calculated Euclidean distance, determine the candidate values of the neighborhood radius eps, and select the optimal eps that makes the density within the cluster stable to adjust the eps of the sliding window; Assign semantic labels to the features extracted by the adjusted sliding window to complete the preliminary classification.

5. The method for detecting cloud system anomalies by automatic matching of time series classification according to claim 4, characterized in that, Performing differential encoding and data compression on the constructed second-level time series and then sending it to the master node includes: Map each semantic label to a binary code as the classification code; Embed the classification code in the header metadata of the second-level time series obtained by performing position transformation, and attach key statistics; Only use the classification code, the attached key statistics, and the representative data points of the header sequence as the transmission data, and perform compression processing on the transmission data and then send it to the master node.

6. The method for abnormal detection of a cloud system with automatic matching of time series classification according to claim 1, characterized in that, The pre-configured anomaly detection algorithms with different parameters include: For periodic data, configure a detection algorithm for seasonal-trend decomposition using LOESS (STL) + kσ rule, and the configured parameters include the period length and the confidence interval threshold, where σ represents the standard deviation, which is used to characterize the degree of dispersion of the time series within the window; For mutant data, adopt the cumulative sum control chart (CUSUM) algorithm, and the configured parameters include the baseline mean and the offset threshold; For low signal-to-noise ratio data, use wavelet transform to suppress the data and detect it through the isolation forest, and the configured parameters include the number of wavelet decomposition levels and the number of trees; For trend data, adopt the LSTM prediction model, and the configured parameters include the sliding window size and the prediction step length.

7. The method for detecting system anomalies in the cloud for automatic matching of time series classification according to claim 6, wherein, Based on the recognition of the header metadata of the obtained secondary time series, the algorithm matching includes: Establish the matching relationship between the label and the algorithm according to the classification code embedded in the header of the secondary time series.

8. The method for detecting cloud system anomalies by automatic matching of time series classification according to claim 1, characterized in that, It also includes: The master node feeds back the anomaly detection results to the slave nodes to record the feedback data at each slave node and optimize the semantic label results.