A data classification method and device based on large-scale network
By using one-dimensional convolutional automatic encoder (1DCAE) and feature selection strategy, the efficiency and accuracy problems of MTS data clustering in large-scale web services are solved, and efficient and robust clustering effect is achieved and the training overhead of the anomaly detection model is reduced.
Patent Information
- Application Number
- CN202210306441.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-03-25
AI Technical Summary
In large-scale web services, it is difficult for prior art to efficiently and accurately cluster multivariate time series (MTS) data of thousands or even hundreds of thousands of system instances, especially in the presence of noise and anomaly data.
One-dimensional convolutional automatic encoder (1DCAE) is used to embed high-dimensional data into low-dimensional data, extract the main features of MTS, and select periodic and representative features through feature selection strategies to reduce clustering overhead and eliminate the impact of noise and anomalies.
Accurate and efficient clustering of the normal mode of the system instance MTS is realized, which reduces the training overhead of the anomaly detection model and improves the robustness of clustering.
Smart Images

Figure CN114861753B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data detection and classification, and in particular to a data classification method and device based on a large-scale network. Background Art
[0002] Web services are becoming increasingly large, often running thousands or even hundreds of thousands of system instances on different containers, virtual machines, or physical machines. The reliability of these system instances is critical to Web services, and abnormal behaviors occurring on system instances may reduce the availability of Web services, affect user experience, and even lead to huge economic losses. In reality, monitoring indicator data is usually recorded to form a multivariate time series (MTS). A series of deep learning-based methods can accurately learn complex patterns in massive MTS data for MTS anomaly detection.
[0003] However, there are a large number of system instances in large-scale Web services (for example, Alibaba and ByteDance have millions of system instances). Training the MTS anomaly detection model for each system instance will consume a lot of computing resources. On the other hand, the complex data patterns in the MTS data of different system instances may be quite different. Training an anomaly detection model for all system instances will reduce the accuracy of anomaly detection for different system instances. Therefore, deploying these MTS anomaly detection methods in large-scale Web services is a very challenging problem.
[0004] Existing methods include Copulas, Mc2PCA, FCFW and TICC, which can cluster MTS data; CTF can cluster data first and then detect anomalies. Copulas considers the relationship between two variables in a single MTS, and performs density-based nonparametric estimation by comparing the distance between two MTSs; Mc2PCA constructs a common projection axis for each cluster, and distributes data to different clusters by calculating the reconstruction error on the corresponding common projection axis; FCFW generates clustering results by comparing the distance between two MTSs based on two distance calculation methods - DTW and SBD; TICC focuses on subsequences in MTS and proposes a model-based clustering method. Each cluster in the TICC algorithm is defined by a correlation network that describes the interdependence between different observations in a typical subsequence in the cluster. CTF is a framework designed for OmniAnomaly to improve training efficiency.
[0005] Copulas is affected by the explosion of dimensionality and has a high computational cost; Mc2PCA only considers the similarity within the cluster and does not consider the similarity between clusters, which may lead to too many clusters; the time complexity of the two algorithms DTW and SBD used by FCFW is very high and cannot be applied to large-scale data; TICC simultaneously segments and clusters MTS data, which consumes a lot of time and computing space and cannot be applied to large-scale data. At the same time, the above four algorithms are designed for ideal smooth data and do not consider the existence of noise and abnormal data in the data collected in real scenarios. These noise and anomalies will greatly affect the clustering effect. In general, the existing clustering methods cannot efficiently and accurately cluster data that is huge in scale (number of system instances, number of indicators, number of time points) and contains noise and anomalies. CTF can only be used in combination with specific anomaly detection algorithms, but not with other anomaly detection algorithms, which has great limitations. Summary of the invention
[0006] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0007] To this end, the purpose of the present invention is to propose a data classification method based on a large-scale network, which uses a one-dimensional convolutional autoencoder (1DCAE) to embed high-dimensional data into low-dimensional data to extract the main features of MTS and embed them into low-dimensional data, which can effectively reduce the clustering overhead and eliminate the influence of noise and anomalies. In addition, an efficient and effective strategy is adopted to select periodic and representative features to prevent certain features from interfering with the MTS clustering effect. The present invention is an efficient and robust solution that can achieve accurate and efficient clustering of the normal mode of the system instance MTS and effectively reduce the training overhead of the anomaly detection model.
[0008] Another object of the present invention is to provide a data classification device based on a large-scale network.
[0009] To achieve the above object, the present invention proposes a data classification method based on a large-scale network, comprising:
[0010] Acquire data to be detected; wherein the data to be detected includes system-level indicators and user-level indicators; perform data preprocessing of smoothing and normalizing the multivariate time series of the data to be detected to obtain preprocessed data; input the preprocessed data into a one-dimensional convolutional autoencoder trained by offline clustering for data compression processing, and perform feature selection using the feature index obtained by the offline clustering, perform distance calculation based on the result of the feature selection to perform online data classification; based on the online data classification, output the online classification result of the data to be detected.
[0011] In addition, the large-scale network-based data classification method according to the above embodiment of the present invention may also have the following additional technical features:
[0012] Furthermore, in one embodiment of the present invention, the system-level indicators include: multiple ones of CPU utilization, memory utilization, disk I / O and network throughput; the user-level indicators include: multiple ones of average response time, error rate and page views.
[0013] Furthermore, in one embodiment of the present invention, the one-dimensional convolutional autoencoder is trained, including: performing the data preprocessing on the multivariate time series of the data to be detected offline to obtain the preprocessed data; using the preprocessed data to train the one-dimensional convolutional autoencoder and compressing the number of time points on each variable of the preprocessed data to obtain a first hidden representation; performing the feature selection on the first hidden representation to obtain a feature index, and performing offline clustering based on the feature index by clustering to obtain a cluster center.
[0014] Furthermore, in one embodiment of the present invention, the feature index obtained by offline clustering is used to perform feature selection, and distance calculation is performed based on the result of the feature selection to perform online data classification, including: using a one-dimensional convolutional autoencoder trained by offline clustering to compress the number of time points on each variable of the preprocessed data to obtain a second hidden representation; using the feature index to perform feature selection on the second hidden representation to obtain a third hidden representation; calculating the distance between the third hidden representation and the cluster center, and selecting the cluster corresponding to the cluster center with the shortest distance as the category for the online data classification.
[0015] Furthermore, in one embodiment of the present invention, the data preprocessing includes:
[0016] Linear interpolation is used to fill in deleted or missing values of the multivariate time series MTS. The baseline of the MTS curve is extracted by the sliding window sliding average algorithm to smooth the MTS curve. Normalization is used in all data to scale each data point to the range of [0,1]. The normalization formula is:
[0017]
[0018] Furthermore, in one embodiment of the present invention, the feature selection includes: deleting non-periodic features, constructing a redundant feature matrix, and deleting redundant features.
[0019] Further, in one embodiment of the present invention, the deleting of the non-periodic features comprises: extracting periodic information using YIN, and obtaining the retained features after deleting the non-periodic features; wherein YIN(z sm)>0 indicates feature z sm There is periodicity, YIN(z sm )=0 indicates feature z sm There is no periodic pattern; the construction of the redundant feature matrix includes: constructing a redundant feature matrix R∈[0,1] M′×M′ , and use the normalized cross-correlation function to calculate whether there is redundancy between the two features; where M' represents the number of features retained after deleting non-periodic features, R ij >0 means there is redundancy between feature i and feature j, R ij =0 There is no redundancy between feature i and feature j; the deleting redundant features includes: defining a set of unassigned features F, F contains the indexes of all M' features, applying the preset feature selection rules from the first rule to the fourth rule to F in sequence, until all features are assigned to the selected feature set SF or the feature set DF is deleted, and all selected features in SF are spliced into z", as the input of the clustering or the classification.
[0020] Furthermore, in one embodiment of the present invention, a hierarchical clustering method is used to cluster z", each piece of data is initialized as a cluster, the inter-cluster distance is iteratively calculated, and the clusters whose inter-cluster distance is lower than the distance threshold are merged until all the inter-cluster distances are greater than the distance threshold.
[0021] Furthermore, in one embodiment of the present invention, the inter-cluster distance is the Euclidean distance:
[0022]
[0023] Here, |*| represents the size of the set and M' is the number of indices in SF.
[0024] The large-scale network-based data classification method of the embodiment of the present invention adopts an efficient and effective strategy to select periodic and representative features to prevent certain features from interfering with the MTS clustering effect. In addition, the method of the present invention is an efficient and robust solution that can accurately and efficiently cluster the normal mode of the system instance MTS and effectively reduce the training overhead of the anomaly detection model.
[0025] To achieve the above object, the present invention proposes a data classification device based on a large-scale network, comprising:
[0026] A data acquisition module, used to acquire the data to be detected; wherein the data to be detected includes system-level indicators and user-level indicators;
[0027] A data processing module, used for performing data preprocessing on the multivariate time series of the data to be detected to smooth and normalize the data to obtain preprocessed data;
[0028] A data classification module, used for inputting the pre-processed data into a one-dimensional convolutional autoencoder trained by offline clustering for data compression processing, performing feature selection using the feature index obtained by the offline clustering, and performing distance calculation based on the result of the feature selection to perform online data classification;
[0029] The result output module is used to output the online classification result of the data to be detected based on the online data classification.
[0030] The large-scale network-based data classification device of the embodiment of the present invention adopts an efficient and effective strategy to select periodic and representative features to prevent certain features from interfering with the MTS clustering effect. In addition, the method of the present invention is an efficient and robust solution that can accurately and efficiently cluster the normal mode of the system instance MTS and effectively reduce the training overhead of the anomaly detection model.
[0031] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0033] Figure 1 is a flow chart of a large-scale network-based data classification method according to an embodiment of the present invention;
[0034] Figure 2 Design an overall structure diagram for the clustering part according to an embodiment of the present invention;
[0035] Figure 3 is an overall schematic diagram of an abnormality detection part according to an embodiment of the present invention;
[0036] Figure 4 1D-CAE model according to an embodiment of the present invention is a schematic diagram showing the structure;
[0037] Figure 5 A schematic diagram of a feature selection process according to an embodiment of the present invention;
[0038] Figure 6 Schematic diagram of the structure of a large-scale network-based data classification device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0040] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0041] The following describes a large-scale network-based data classification method and device according to an embodiment of the present invention with reference to the accompanying drawings.
[0042] Figure 1 The figure is a flow chart of a large-scale network-based data classification method according to an embodiment of the present invention.
[0043] like Figure 1 As shown, the method includes but is not limited to the following steps:
[0044] S1, obtaining data to be detected; wherein the data to be detected includes system-level indicators and user-level indicators;
[0045] S2, performing data preprocessing on the multivariate time series of the data to be detected by smoothing and normalizing to obtain preprocessed data;
[0046] S3, input the preprocessed data into the one-dimensional convolutional autoencoder trained by offline clustering for data compression processing, and use the feature index obtained by offline clustering to perform feature selection, and perform distance calculation based on the result of feature selection to perform online data classification;
[0047] S4: Based on the online data classification, output the online classification result of the data to be detected.
[0048] It is understandable that in order to proactively detect abnormal behaviors of system instances and promptly mitigate system failures, Web service operators configure different types of system-level metrics (e.g., CPU utilization, memory utilization, disk I / O, network throughput) and user-level metrics (e.g., average response time, error rate, page views) in the system and continuously collect monitoring data at predetermined time intervals.
[0049] Specifically, the acquired monitoring data is used for subsequent data clustering and classification.
[0050] like Figure 2 As shown in the figure, the clustering design of the present invention consists of two main parts: offline clustering and online classification. The overall structure is as follows: Figure 2 As shown, the solid line represents offline clustering and the dotted line represents online classification.
[0051] Offline clustering is divided into four process stages. The first stage is data preprocessing, which first smoothes and normalizes the MTS data; the second stage trains 1D-CAE and compresses the number of time points on each variable to obtain the hidden representation z; the third stage performs feature selection on z to reduce the number of variables in each data to obtain z'; the last stage uses a hierarchical clustering method to cluster the data.
[0052] Online classification is also divided into four stages. The first stage is the same data preprocessing as the first stage of offline clustering, which smoothes and normalizes the data. The second stage uses the 1D-CAE encoder trained in the second stage of offline clustering to compress the number of time points on each variable to obtain the hidden representation z. The third stage uses the feature (variable) index obtained in the third stage of offline clustering to perform feature selection on z to obtain z'. The last stage calculates the distance between z' and the cluster center obtained in the fourth stage of offline clustering, and selects the cluster corresponding to the cluster center with the shortest distance as the data category.
[0053] Furthermore, the overall design of the anomaly detection part is as follows Figure 3 As shown:
[0054] The anomaly detection part is divided into two parts: offline training of anomaly detection model and online detection of data anomalies.
[0055] The present invention uniformly uses x smt represents the data of the mth variable of the sth system instance at the tth time point; the sth system instance x s It is an M×T matrix, where there are M variables and T time points of monitoring data in the system instance.
[0056] The following is a detailed description of the embodiments of the invention in conjunction with the accompanying drawings:
[0057] Data preprocessing. There are usually noise, anomalies and missing values in MTS that will significantly affect the shape of the data, and their negative impact must be minimized. Since extreme values are usually more likely to become anomalies, the present invention uses the method of deleting the top 5% of data that deviates from the mean to deal with extreme values; and in real production situations, there may be errors in the data collection process that cause some missing values in the data. The present invention uses linear interpolation to fill in deleted or missing values; in order to deal with noise, the present invention extracts the baseline of the MTS curve through a sliding window sliding average algorithm to smooth the MTS curve. Finally, in order to deal with the amplitude differences between different data, normalization is used in all data to scale each data point to the range of [0,1]. The specific normalization formula is:
[0058]
[0059] The same data preprocessing steps are used for both offline clustering and online classification.
[0060] 1D-CAE compresses data. In order to reduce the impact of high data dimensions on clustering efficiency, such as Figure 4 It is shown that the present invention uses 1D-CAE and reconstruction loss function for model training to effectively reduce the data dimension and capture the nonlinear characteristics of the data.
[0061] It can be understood that the autoencoder (AE) includes two basic units: an encoder and a decoder. The encoder compresses the input into a latent space representation, and the decoder uses the latent space representation to reconstruct the input data. AE can optimize model parameters by minimizing the difference between input and output (reconstruction loss). The convolutional autoencoder (CAE) uses a convolutional neural network (CNN) encoder and decoder. In the present invention, 1D-CAE is used for feature extraction and dimensionality reduction. The convolutional encoder can learn the normal pattern of the input data and ignore noise and anomalies. In the encoder of the present invention, each variable in the MTS is input into a different convolutional neural network, and M corresponding features can be obtained. The convolutional decoder consists of M one-dimensional deconvolutional neural networks with independent parameters. The encoder output is the compressed feature z, which is an M×T′ matrix, where T′ is the dimension of each variable; the output of the decoder is the reconstruction of the original data The size is consistent with the input data x. The figure below shows a schematic diagram of the 1D-CAE model. When the input MTS has three variables, the encoder of 1D-CAE is composed of three convolutional neural networks, and the decoder is composed of three one-dimensional deconvolutional neural networks.
[0062] Furthermore, the mean square error is used as the loss function in the offline clustering process, by minimizing the input data x and the output data The loss between , and the 1D-CAE model is continuously updated. Finally, the encoder structure and parameters of 1D-CAE are saved and the data z is obtained as the input of the next stage. Online classification uses the encoder of 1D-CAE saved by offline clustering to obtain the data z as the input of the next stage.
[0063] Feature selection. This paper implements a robust general feature selection method to reduce the number of features in the metric dimension and improve clustering performance. The feature selection process includes three steps: deleting non-periodic features, building a redundant feature matrix, and deleting redundant features. Figure 5 As shown:
[0064] 1) Remove non-periodic features:
[0065] First, use YIN to extract periodic information. YIN(z sm )>0 indicates feature z sm There is periodicity, and YIN(z sm )=0 indicates feature z sm There is no obvious periodic pattern. The indexes corresponding to the features that are non-periodic in most system instances will be deleted, as shown in Algorithm 1. The features retained after deleting the non-periodic features are denoted by z′.
[0066] Algorithm 1
[0067]
[0068] 2) Construct redundant feature matrix:
[0069] Construct the redundant feature matrix R∈[0,1] M′×M′ (M' represents the number of features retained after removing non-periodic features), where R ij >0 means there is redundancy between feature i and feature j, R ij = 0 There is no redundancy between feature i and feature j. Use the normalized cross-correlation function (NCC) to calculate whether there is redundancy between two features. The specific construction scheme is shown in Algorithm 2.
[0070] Algorithm 2
[0071]
[0072]
[0073] 3) Remove redundant features:
[0074] The present invention applies feature selection rules to utilize redundant features in the redundant matrix R. First, a set of unassigned features F is defined, and F contains the indexes of all M' features. Then, the following feature selection rules are sequentially applied to F from rule 1 to rule 4, until all features are assigned to the selected feature set SF or the deleted feature set DF. Finally, the present invention concatenates all selected features in SF into z", as input for the clustering or classification step.
[0075] Rule 1: If R i is completely unrelated to other rows in R, i.e. R i Contains only zeros:
[0076] (a) Add i to the selected feature set SF: SF = SF∪{i};
[0077] (b) Delete i from F and remove the items related to i in R;
[0078] Rule 2: If R i All features in R are correlated with other rows and there is at least one feature in R that is not fully correlated with other features (that is, R has a value of 0 on the off-diagonal line):
[0079] (a) Add i to the deletion feature set DF: DF = DF∪{i};
[0080] (b) Delete i from F and remove the items related to i in R;
[0081] Rule 3: If all features in F are correlated with each other (i.e. R contains only non-zero off-diagonal values):
[0082] (a) Select feature i with the smallest correlation with the features contained in Sf;
[0083] (b) Add i to the selected feature set SF: SF = SF∪{i};
[0084] (c) Delete i from F and remove the items related to i from R;
[0085] (d) Move the remaining features in F to the deleted feature set DF: DF = DF∪{i} and terminate.
[0086] Rule 4: If neither Rule 2 nor Rule 3 applies:
[0087] (a) Select feature i with the smallest correlation with the features contained in F;
[0088] (b) Definition is the feature related to i in F, and then select the feature j∈S(i) with the greatest correlation with the feature contained in SF;
[0089] (c) Add i to the selection feature set SF: SF = SF∪{i}, add j to the deletion feature set DF: DF = DF∪{j};
[0090] (d) Delete i, j from F and remove the items related to i, j in R;
[0091] In Rules 3(a), 4(a), and 4(b), the following formula is used to select features: Where R is the constructed redundant feature moment.
[0092] During offline clustering, feature selection is performed, the selected feature set SF is obtained and saved, and the selected features are concatenated into z” as the input for the next stage. Online classification uses the selected feature set SF saved by offline clustering to obtain z” of data as the input for the next stage.
[0093] Clustering and classification. In the offline clustering stage, the present invention uses a hierarchical clustering scheme to cluster z'. First, each data is initialized as a cluster, and then the Euclidean distance between clusters is iteratively calculated. |*| represents the size of the set, and M” is the number of indices in SF and the number of features retained) and merging clusters whose inter-cluster distances are lower than the distance threshold until the distances between all clusters are greater than the distance threshold. All cluster center data are retained in the offline clustering stage.
[0094] In the online classification phase, the Euclidean distance between the features extracted from the data and the data of all cluster centers retained in the offline clustering phase is calculated, and then the cluster category corresponding to the nearest cluster center is selected as the category of the data. It is particularly noted that if the distance between the nearest cluster center and the data is also greater than the distance threshold, the data will not be classified, but reported to the task executor as abnormal data, which also enhances the robustness of the present invention.
[0095] Anomaly detection. In the anomaly detection offline training part, the present invention trains an anomaly detection model (which can be any existing anomaly detection model) for each cluster center obtained by clustering, and uses a training scheme consistent with the anomaly detection model used for training.
[0096] The present invention collects real-time online data, and uses an anomaly detection model corresponding to the data category obtained by cluster classification to perform anomaly detection on the online data.
[0097] Furthermore, after investigating thousands of real-world system instances, the present invention can use clustering to automatically group system instances into different clusters, where system instances in each cluster have similar patterns. Therefore, an MTS anomaly detection model can be trained for each cluster instead of each system instance, which can significantly reduce training overhead because the number of clusters is much smaller than that of system instances.
[0098] As an implementation method, the existing traditional K-Means algorithm or Copulas, Mc2PCA, FCFW and TICC can be used as alternatives to the present invention, but the effect and efficiency are not as ideal as those of the present invention.
[0099] As an implementation method, the clustering method used in the present invention is hierarchical clustering, which can be replaced by DBSCAN; the Euclidean distance used when calculating the distance between data can be replaced by Manhattan distance, SBD, etc.
[0100] Furthermore, as an implementation method, any computer language can be used, and there is no special requirement for the software and hardware environment.
[0101] Preferably, the present invention is implemented using Python 3.8 as the computer language, Tensorflow 2.2 as the software environment, and a 16C32T Intel(R) Xeon(R) Gold 5218 CPU @ 2.30 GHz and 192 GB of RAM as the hardware environment, which can be recommended.
[0102] According to the large-scale network-based data classification method of the embodiment of the present invention, a one-dimensional convolutional autoencoder (1DCAE) is used to embed high-dimensional data into low-dimensional data to extract the main features of MTS and embed them into low-dimensional data, which can effectively reduce the clustering overhead and eliminate the influence of noise and anomalies. In addition, an efficient and effective strategy is adopted to select periodic and representative features to prevent certain features from interfering with the MTS clustering effect. The present invention is an efficient and robust solution that can achieve accurate and efficient clustering of the normal mode of the system instance MTS and effectively reduce the training overhead of the anomaly detection model.
[0103] In order to implement the above embodiment, Figure 6 As shown, this embodiment also provides a large-scale network-based data classification device 10, which includes: a data acquisition module 100, a data processing module 200, a data classification module 300 and a result output module 400.
[0104] The data acquisition module 100 is used to acquire the data to be detected; wherein the data to be detected includes system-level indicators and user-level indicators;
[0105] The data processing module 200 is used for performing data preprocessing on the multivariate time series of the detected data to obtain preprocessed data;
[0106] The data classification module 300 is used to input the pre-processed data into the one-dimensional convolutional autoencoder trained by offline clustering for data compression processing, and perform feature selection using the feature index obtained by offline clustering, and perform distance calculation based on the result of feature selection to perform online data classification;
[0107] The result output module 400 is used to output the online classification result of the data to be detected based on the online data classification.
[0108] According to the large-scale network-based data classification device of the embodiment of the present invention, a one-dimensional convolutional autoencoder (1DCAE) is used to embed high-dimensional data into low-dimensional data to extract the main features of MTS and embed them into low-dimensional data, which can effectively reduce the clustering overhead and eliminate the influence of noise and anomalies. In addition, an efficient and effective strategy is adopted to select periodic and representative features to prevent certain features from interfering with the MTS clustering effect. The present invention is an efficient and robust solution that can achieve accurate and efficient clustering of the normal mode of the system instance MTS and effectively reduce the training overhead of the anomaly detection model.
[0109] It should be noted that the aforementioned explanation of the embodiment of the data classification method based on a large-scale network is also applicable to the data classification device based on a large-scale network of this embodiment, and will not be repeated here.
[0110] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0111] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0112] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A data classification method based on a large-scale network, characterized in that: The following steps are involved: Acquire the data to be detected; wherein the data to be detected includes system-level indicators and user-level indicators; Performing data preprocessing by smoothing and normalizing the multivariate time series of the data to be detected to obtain preprocessed data; The preprocessed data is input into a one-dimensional convolutional autoencoder trained by offline clustering for data compression processing, and feature selection is performed using the feature index obtained by the offline clustering, and distance calculation is performed according to the result of the feature selection to perform online data classification; Based on the online data classification, outputting the online classification result of the data to be detected; The feature selection includes: deleting non-periodic features, constructing a redundant feature matrix, and deleting redundant features; The deleting of non-periodic features includes: extracting periodic information using YIN, and obtaining retained features after deleting the non-periodic features; wherein, Representation characteristics There is periodicity, Representation characteristics There is no cyclical pattern; The constructing of the redundant feature matrix includes: constructing a redundant feature matrix , and use the normalized cross-correlation function to calculate whether there is redundancy between the two features; where, represents the number of features retained after deleting non-periodic features, Representation characteristics and Features There is redundancy between feature and Features There is no redundancy between them; Deleting redundant features includes: defining a set of unassigned features , Contains the index of all 𝑀' features, and applies the preset feature selection rules from the first rule to the fourth rule in sequence. , until all features are assigned to the selection feature set or delete a feature set ,Will All selected features in are concatenated into , as input for the clustering or the classification.
2. The method according to claim 1, characterized in that The system-level indicators include: multiple ones of CPU utilization, memory utilization, disk I / O and network throughput; the user-level indicators include: multiple ones of average response time, error rate and number of page views.
3. The method according to claim 1, characterized in that The one-dimensional convolutional autoencoder is trained, comprising: Offline preprocessing of the multivariate time series of the data to be detected is performed to obtain the preprocessed data; Using the preprocessed data to train a one-dimensional convolutional autoencoder and compressing the number of time points on each variable of the preprocessed data to obtain a first hidden representation; The feature selection is performed on the first hidden representation to obtain a feature index, and offline clustering is performed based on the feature index in a clustering manner to obtain a cluster center.
4. The method according to claim 3, characterized in that The performing of feature selection using the feature index obtained by the offline clustering, and performing distance calculation according to the result of the feature selection to perform online data classification, includes: Using a one-dimensional convolutional autoencoder trained by offline clustering to compress the number of time points on each variable of the preprocessed data, to obtain a second hidden representation; Using the feature index, feature selection is performed on the second hidden representation to obtain a third hidden representation; the distance between the third hidden representation and the cluster center is calculated, and the cluster corresponding to the cluster center with the shortest distance is selected as the category of the online data classification.
5. The method according to claim 1, characterized in that The data preprocessing includes: Linear interpolation is used to fill in deleted or missing values of the multivariate time series MTS. The baseline of the MTS curve is extracted by the sliding window sliding average algorithm to smooth the MTS curve. Normalization is used in all data to scale each data point to the range of [0,1]. The normalization formula is: 。 6. The method according to claim 1, characterized in that Using hierarchical clustering Perform clustering, initialize each data as a cluster, iteratively calculate the inter-cluster distance, and merge the clusters whose inter-cluster distance is lower than the distance threshold until all the inter-cluster distances are greater than the distance threshold.
7. The method according to claim 6, characterized in that The inter-cluster distance is the Euclidean distance: 2 Among them, |∗| represents the size of the set, 𝑀′′ is The number of indexes in .
8. A data classification device based on a large-scale network, characterized in that: include: A data acquisition module, used to acquire the data to be detected; wherein the data to be detected includes system-level indicators and user-level indicators; A data processing module, used for performing data preprocessing on the multivariate time series of the data to be detected to smooth and normalize the data to obtain preprocessed data; A data classification module, used for inputting the pre-processed data into a one-dimensional convolutional autoencoder trained by offline clustering for data compression processing, performing feature selection using the feature index obtained by the offline clustering, and performing distance calculation based on the result of the feature selection to perform online data classification; A result output module, used for outputting the online classification result of the data to be detected based on the online data classification; The feature selection includes: deleting non-periodic features, constructing a redundant feature matrix, and deleting redundant features; The deleting of non-periodic features includes: extracting periodic information using YIN, and obtaining retained features after deleting the non-periodic features; wherein, Representation characteristics There is periodicity, Representation characteristics There is no cyclical pattern; The constructing of the redundant feature matrix includes: constructing a redundant feature matrix , and use the normalized cross-correlation function to calculate whether there is redundancy between the two features; where, represents the number of features retained after deleting non-periodic features, Representation characteristics and Features There is redundancy between feature and Features There is no redundancy between them; Deleting redundant features includes: defining a set of unassigned features , Contains the index of all 𝑀' features, and applies the preset feature selection rules from the first rule to the fourth rule in sequence. , until all features are assigned to the selection feature set or delete a feature set ,Will All selected features in are concatenated into , as input for the clustering or the classification.
Citation Information
Patent Citations
Multivariable time series data clustering method
CN111488924A
Multivariable time series data classification method based on FCN
CN112465054A