Intelligent abnormity diagnosis method and system based on distributed cloud database
By building a hierarchical dependency graph and using Triplet Loss loss function to train the twin differential network, the problem of low accuracy of database abnormal diagnosis in the existing technology is solved, and higher abnormal diagnosis accuracy and cross-cluster generalization performance are achieved.
Patent Information
- Application Number
- CN202510042032.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing intelligent database anomaly diagnosis methods are difficult to capture complex anomaly patterns, resulting in a small amount of training data in database anomaly diagnosis and poor generalization performance across clusters, resulting in a low accuracy rate of abnormal diagnosis.
By building a hierarchical dependency graph of KPI time series data based on the database node cluster, obtain the training samples formed by pairing abnormal samples and normal samples in the corresponding cluster, and use the Triplet Loss loss function to train the twin differential network to improve cross-cluster generalization and alleviate the overfitting problem caused by small data volume.
It improves the accuracy of abnormal diagnosis of distributed databases, enhances feature extraction capabilities, improves generalization performance across clusters, and solves the overfitting problem caused by small data volume.
Smart Images

Figure CN119938382A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of cloud database anomaly diagnosis, and more specifically, to an abnormality intelligent diagnosis method and system based on a distributed cloud database. Background Art
[0002] With the explosive growth of data and the growth of business scale, database systems have evolved from stand-alone architecture to distributed cloud architecture, resulting in increasingly complex underlying designs. In recent years, distributed cloud database systems have developed rapidly and have become an important part of infrastructure services, which are mainly provided by cloud providers. Due to the complexity of the architecture, distributed cloud database systems face various complex performance anomalies. For example, in distributed nodes, problems such as I / O-intensive task overload and data skew can cause serious performance degradation. If these problems are not discovered and resolved in a timely manner, they may even cause service crashes, causing unpredictable losses to users. This is especially critical in cloud database scenarios, as service providers must work hard to ensure service quality.
[0003] In order to detect anomalies, it is usually necessary to analyze the time series of key performance indicators (KPIs) to determine whether the system is abnormal. However, anomaly detection alone is not enough. In order to effectively eliminate anomalies, it is crucial to identify potential root causes. For most database administrators, analyzing the root causes of anomalies is very challenging. Modern databases are usually equipped with monitoring systems that collect KPI data describing the system status at fixed time intervals. Over time, a large amount of data is collected, and manual analysis of this data is very time-consuming and labor-intensive. Therefore, automated diagnostic methods have emerged, some of which are used to locate root cause KPIs or structured query languages (SQL), and others are used to identify the type of root cause.
[0004] In the existing database anomaly diagnosis research, DBSherlock and iSQUAD are statistical anomaly diagnosis methods designed for database systems. DBsherlock obtains horizontal anomaly features of KPIs by mining predicates. iSQUAD extracts 4 predefined trend anomaly patterns through statistical tests. Dundjerski et al. proposed a rule-based end-to-end automatic anomaly diagnosis system for the Azure SQL cloud database system. All rules are manually designed for specific database systems, which requires a lot of manual work. The anomaly features they extract are limited and cannot cope with diverse anomaly patterns. And most of them rely on expert knowledge, summarize the prior features or rules of anomalies, and cannot fully cover various patterns of anomalies.
[0005] In recent years, research on anomaly diagnosis based on deep learning has also been proposed. For example, an anomaly diagnosis framework for high-performance computing systems based on active learning alleviates the problem of small data volume to a certain extent through active learning, but it does not consider the generalization problem under different configuration systems. Another example is a responsive and explainable online service system fault location method based on deep learning, which designs a unified model to extract the features of each component of the system and designs corresponding diagnostic interpretation methods. It is applicable to various fault scenarios of the system, but does not consider the problem of model overfitting when the amount of data is small. The above two methods are not designed for database systems and do not consider the characteristics of database anomalies. Summary of the invention
[0006] In view of the defects of the prior art, the purpose of this application is to provide an abnormality intelligent diagnosis method and system based on a distributed cloud database, aiming to solve the problem of low abnormality diagnosis accuracy caused by the small amount of training data and poor cross-cluster generalization performance when applying deep learning technology to database anomaly diagnosis problems, due to the difficulty of existing database anomaly intelligent diagnosis methods in capturing complex abnormal patterns.
[0007] To achieve the above objectives, in a first aspect, the present application provides an abnormal intelligent diagnosis method based on a distributed cloud database, comprising: Build a hierarchical dependency graph based on the KPI time series data of the database node cluster; Acquire training samples, where the training samples include abnormal samples and samples formed by pairing normal samples in a cluster corresponding to the abnormal samples; Based on the training samples and the hierarchical dependency graph, the twin difference network is trained using a Triplet Loss loss function to obtain a trained twin difference network; Based on the trained twin difference network, the abnormal root cause type of the database node cluster to be tested is determined.
[0008] This application uses the Triplet Loss loss function to train the twin difference network by taking both abnormal data and the normal data of its corresponding cluster as training samples, combining the hierarchical dependency graph, to improve the cross-cluster generalization while alleviating the overfitting problem caused by the small amount of data. The trained twin difference network is used for intelligent diagnosis of database anomalies to improve the accuracy of anomaly diagnosis.
[0009] According to an abnormal intelligent diagnosis method based on a distributed cloud database provided by the present application, the key performance indicator KPI time series data of the database node cluster is used to construct a hierarchical dependency graph, including: Use the Isolation Forest algorithm to clean the normal KPI time series data of the database node cluster; Based on the cleaned normal KPI time series data and the abnormal KPI time series of the database node cluster, a KPI dependency graph is constructed; The KPI dependency graph is expanded into a hierarchical dependency graph.
[0010] According to an abnormal intelligent diagnosis method based on a distributed cloud database provided by the present application, based on the training sample and the hierarchical dependency graph, the twin difference network is trained using the Triplet Loss loss function to obtain a trained twin difference network, including: Constructing a twin difference network, the twin difference network is composed of three shared weight difference networks, including a multi-scale trend convolution module and a hierarchical graph convolution module, wherein the multi-scale trend convolution module is used to extract short-term and long-term trends and horizontal features from the KPI time series data, and the hierarchical graph convolution module is used to mine the correlation of the KPI time series data based on the hierarchical dependency graph; The constructed twin difference network is trained based on the training samples to obtain a trained twin difference network.
[0011] Different input structures are designed in the twin difference network constructed in this application, and based on the abnormal characteristics of the database system, multi-scale trend convolution modules and hierarchical graph convolution modules are designed to enhance feature extraction capabilities.
[0012] According to an abnormal intelligent diagnosis method based on a distributed cloud database provided by the present application, the abnormal root cause type of the database node cluster to be tested is determined based on the trained twin difference network, including: Obtaining a difference sample, where the difference sample is a pair formed by an abnormal sample and a normal sample in the database node cluster to be tested; The difference sample is input into the trained twin difference network to determine the abnormal root cause type of the database node cluster to be tested.
[0013] According to an abnormal intelligent diagnosis method based on a distributed cloud database provided by the present application, the difference sample is input into the trained twin difference network to determine the abnormal root cause type of the database node cluster to be tested, including: Inputting the difference sample into the trained twin difference network to generate a difference representation corresponding to the difference sample; Based on the difference representation and the prototype difference representation of the training sample, the abnormal root cause type of the database node cluster to be tested is determined.
[0014] According to an abnormal intelligent diagnosis method based on a distributed cloud database provided by the present application, the method also includes: Transforming the difference sample and inputting it into the trained twin difference network to obtain the difference representation corresponding to the transformed sample; Obtaining a symptom KPI set based on the difference representation corresponding to the transformed sample and the deviation of the difference representation corresponding to the difference sample; The Personalized PageRank Vector corresponding to each KPI in the symptom KPI set is calculated, and the root cause KPI causing the abnormality of the database node cluster to be tested is determined based on the Personalized PageRank Vector.
[0015] In a second aspect, the present application provides an abnormal intelligent diagnosis system based on a distributed cloud database, comprising: A construction module is used to build a hierarchical dependency graph based on the key performance indicator (KPI) time series data of the database node cluster; An acquisition module, used to acquire training samples, wherein the training samples include abnormal samples and samples formed by pairing normal samples in a cluster corresponding to the abnormal samples; A training module, used to train the twin difference network based on the training samples and the hierarchical dependency graph using a Triplet Loss loss function to obtain a trained twin difference network; A determination module is used to determine the abnormal root cause type of the database node cluster to be tested based on the trained twin difference network.
[0016] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the abnormality intelligent diagnosis method based on a distributed cloud database described in the first aspect or any possible implementation method of the first aspect.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the abnormal intelligent diagnosis method based on a distributed cloud database described in the first aspect or any possible implementation of the first aspect.
[0018] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the abnormal intelligent diagnosis method based on a distributed cloud database described in the first aspect or any possible implementation of the first aspect.
[0019] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0020] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art: (1) By taking both abnormal data and the normal data of its corresponding cluster as training samples, combined with the hierarchical dependency graph, and using the Triplet Loss loss function to train the twin difference network, we can improve the cross-cluster generalization while alleviating the overfitting problem caused by the small amount of data. The trained twin difference network is used for intelligent diagnosis of database anomalies, which can improve the accuracy of distributed database anomaly diagnosis.
[0021] (2) Different input structures are designed in the constructed twin difference network. In view of the abnormal characteristics of the database system, a multi-scale trend convolution module and a hierarchical graph convolution module are designed to enhance the feature extraction capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 It is a flowchart of an abnormal intelligent diagnosis method based on a distributed cloud database provided in an embodiment of the present application; Figure 2 It is a framework diagram of an abnormal intelligent diagnosis method based on a distributed cloud database provided in an embodiment of the present application; Figure 3 It is a schematic diagram of the structure of the twin difference network provided in the embodiment of the present application; Figure 4 It is a schematic diagram of the model training process provided in the embodiment of the present application; Figure 5 is a schematic diagram of a process for a diagnostic test sample provided in an embodiment of the present application; Figure 6 It is a structural diagram of an abnormal intelligent diagnosis system based on a distributed cloud database provided in an embodiment of the present application; Figure 7 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0025] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0026] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0027] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0028] Next, combine Figure 1-Figure 5 The abnormal intelligent diagnosis method based on distributed cloud database provided in the embodiment of the present application is introduced.
[0029] Figure 1 is a flow chart of an abnormal intelligent diagnosis method based on a distributed cloud database provided in an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: Step 100, constructing a hierarchical dependency graph based on the key performance indicator KPI time series data of the database node cluster; Optionally, the KPI time series data may include time series data such as central processing unit (CPU) utilization, input / output (I / O) utilization, memory usage, and the number of active sessions.
[0030] Optionally, the Fisher-Z Test can be used to test the independence between KPI time series data and construct a hierarchical dependency graph.
[0031] A hierarchical dependency graph is a data structure that describes the correlation between nodes. It can be composed of two layers: the lower layer is the KPI dependency graph, and the higher layer is the KPI Group dependency graph. The KPI dependency graph reflects the correlation between KPIs, and the KPI Group is the grouping of KPIs according to the system components they describe.
[0032] This application first obtains the KPI time series data of the database node cluster, and then builds a hierarchical dependency graph based on the data.
[0033] Step 110, obtaining training samples, where the training samples include abnormal samples and samples formed by pairing normal samples in the cluster corresponding to the abnormal samples; In order to learn the difference between normal samples and abnormal samples, abnormal samples and normal samples can be paired to form difference samples to facilitate subsequent difference learning.
[0034] Specifically, to improve efficiency, firstly, for each cluster, we need to sample from its normal samples The samples are taken as the normal sample representatives of the cluster, and then the abnormal samples are paired with the normal sample representatives of their corresponding clusters to form difference samples. .
[0035] Optionally, when sampling normal samples, a clustering-based strategy can be used to maintain the diversity of normal sample patterns as much as possible during the sampling process in order to learn consistent difference representations about normal patterns. For each cluster, the K-medoids clustering algorithm is used on the normal samples to obtain The algorithm divides the sample clusters into several clusters and then samples the samples in the cluster center. The reason is that samples in the same cluster often share similar patterns, resulting in information redundancy, while samples from different clusters are more likely to show different patterns.
[0036] In the K-medoids algorithm, the channel-independent Euclidean distance is used to measure the distance between normal samples, emphasizing the contribution of single channel distance to the overall distance, which is defined as: ,in represents the time window length, Indicates the KPI quantity.
[0037] Optionally, when sampling normal samples, a random strategy can be used to randomly select, for each cluster, non-repetitive samples from the normal samples of the cluster. A normal sample.
[0038] Step 120, based on the training samples and the hierarchical dependency graph, the twin difference network is trained using the Triplet Loss loss function to obtain a trained twin difference network; This application constructs a twin difference network, which consists of three shared weight difference networks. Through the obtained training samples and hierarchical dependency graph, Triplet Loss is used for training, and the difference representation learned by the difference network is mapped to a Euclidean space, in which the representation distance of similar differences is closer, and the representation distance of different differences is farther.
[0039] Using the Triple Loss loss function can alleviate the problem of small sample size.
[0040] Step 130: Determine the root cause type KPI of the database node cluster to be tested based on the trained twin difference network.
[0041] After the twin difference network is trained, only one branch of the three shared weight difference networks of the model is used. After constructing the hierarchical dependency graph, the difference samples consisting of the abnormal samples of the abnormal database node cluster to be tested and the normal samples of the cluster where the abnormality occurs are input into the trained twin difference network to obtain the difference representation of its output, based on which the root cause type of the abnormality can be determined.
[0042] Figure 2 is a framework diagram of an abnormal intelligent diagnosis method based on a distributed cloud database provided in an embodiment of the present application, such as Figure 2 As shown, in one embodiment of the present application, the KPI data of each cluster is regularly collected by the monitoring system through the data collection module, and is transmitted to the data warehouse for storage through the message queue. The method is implemented in two stages: offline training and online diagnosis. In the offline training stage, historical abnormal cases and normal data are first retrieved from the data warehouse and sent to the data processing module. In the data processing module, the normal data of each cluster is cleaned, and a hierarchical dependency graph is constructed based on the data through correlation analysis, and then a training sample is constructed to train the twin difference network. In the online diagnosis stage, the online diagnosis process is triggered when an abnormality is detected. First, the KPI data of the abnormal period is taken out from the data warehouse and processed by the data processing module. The processed data is input into the trained twin difference network, and the root cause type is determined according to the reasoning result. Finally, the root cause KPI that causes the abnormality is located by data transformation and importance analysis methods.
[0043] The present application provides an abnormality intelligent diagnosis method based on a distributed cloud database. By taking the abnormal data and the normal data of its corresponding cluster as training samples, combining the hierarchical dependency graph, and using the Triplet Loss loss function to train the twin difference network, it improves the cross-cluster generalization while alleviating the overfitting problem caused by the small amount of data. The trained twin difference network is used for intelligent diagnosis of database anomalies, which can improve the accuracy of distributed database anomaly diagnosis.
[0044] In some embodiments, step 100 specifically includes: Step 1001, clean the normal KPI time series data of the database node cluster by using the isolation forest algorithm; Step 1002: construct a KPI dependency graph based on the cleaned normal KPI time series data and the abnormal KPI time series of the database node cluster; Step 1003: Expand the KPI dependency graph into a hierarchical dependency graph.
[0045] In order to robustly extract the difference between normal and abnormal samples, the purity of normal samples must be ensured to avoid contamination of normal patterns by outliers.
[0046] In one embodiment of the present application, a classic and efficient anomaly detection algorithm, Isolation Forest, is used to identify and mark anomalies in the normal KPI time series data of the cluster, and then compare the anomalies with the two sides of the anomaly. The adjacent points are deleted together, the remaining time series is divided into multiple segments, and the sliding window method is used to extract the window with the same length as the abnormal window as the normal window of the cluster.
[0047] Optionally, to reduce redundancy between windows, the sliding step size is set to .
[0048] In order to build a hierarchical dependency graph, first build a KPI dependency graph and put all KPIs The independence of all KPI pairs is then tested and edges are added between KPI nodes that show correlation. Specifically, the widely used Fisher-Z Test is used to test the independence of KPIs. The Fisher-Z Test is constructed by defining a test statistic based on the Pearson correlation coefficient. To test independence. , is the correlation coefficient, is the window length. When the two KPIs are independent It obeys the standard normal distribution, and the independence between KPIs is determined by testing this null hypothesis.
[0049] After constructing the KPI dependency graph, expand it into a hierarchical dependency graph. The hierarchical dependency graph consists of two layers: the lower layer is the KPI dependency graph, and the higher layer is the KPI Group dependency graph. KPI Group is a grouping of KPIs according to the system components they describe. In the KPIGroup dependency graph, KPI Group is a node, and the edges are summarized from the KPI dependency graph. Specifically, for any two KPI Group and , if there is and And there is an edge between them, then in the KPI Group dependency graph and There are edges between them.
[0050] In some embodiments, step 120 specifically includes: Step 1201, constructing a twin difference network, which consists of three shared weight difference networks, including a multi-scale trend convolution module and a hierarchical graph convolution module, wherein the multi-scale trend convolution module is used to extract short-term and long-term trends and horizontal features from the KPI time series data, and the hierarchical graph convolution module is used to mine the correlation of the KPI time series data based on the hierarchical dependency graph; Step 1202: Train the constructed twin difference network based on the training samples to obtain a trained twin difference network.
[0051] In the twin difference network, a multi-scale trend convolution module and a hierarchical graph convolution module are designed to better extract features.
[0052] Figure 3 is a schematic diagram of the structure of the twin difference network provided in the embodiment of the present application, such as Figure 3 As shown in Figure 1, the multi-scale trend convolution module is used to extract short-term and long-term trends and horizontal features from the KPI time series data. In the database scenario, anomalies in KPIs are usually manifested as sudden spikes / drops or sustained growth / decreases. In addition, the continued high or low levels of KPIs may also indicate the presence of anomalies. The multi-scale trend convolution module explicitly extracts short-term and long-term trends and horizontal features from the time series. Specifically, for the time series , the trend convolution uses a sliding window to divide the sequence into pieces of size For each window, linear regression is applied to calculate the regression coefficients and the intercept .
[0053]
[0054]
[0055] in and can be regarded as a trend transformation and identity change of the input. Then, the kernel size is The one-dimensional convolutional layer is applied to and To extract further features.
[0056]
[0057]
[0058] Then the output is max-pooled in the time dimension. The scale is The trend convolution is Map the input to . Use multiple trend convolutions of different scales to capture different features, where the scale parameter exist In general, the multi-scale convolution module maps the input to .
[0059]
[0060]
[0061] The hierarchical graph convolution module is used to mine the correlation of KPI time series data based on the hierarchical dependency graph. In order to use the pre-built hierarchical dependency graph to explicitly mine the correlation of KPIs, this method extends the graph convolution to a hierarchical graph convolution to learn the multi-granularity correlation information of KPIs. The hierarchical graph convolution consists of two layers of graph convolution, taking the output of the multi-scale trend convolution module and the hierarchical dependency graph as input. First, a nonlinear transformation is applied to the output of the multi-scale trend convolution module. The underlying graph convolution applies a graph convolution operation to the KPI dependency graph to aggregate node features and generate .
[0062]
[0063]
[0064] in, Hehe is the parameter in the nonlinear transformation, is the adjacency matrix of the KPI dependency graph, and It is The weights and constants of the layer.
[0065] Next, we aggregate KPI node features by grouping to generate corresponding KPI Group features. Aggregate each The upper-layer graph convolution applies graph convolution operations on the KPI Group dependency graph to aggregate the KPI Group node features and generate .
[0066]
[0067]
[0068] in and is the parameter of the aggregation transformation, is the adjacency matrix of the KPI Group dependency graph, and For the Layer weights and constants. Graph Convolutional Layer and The number depends on the complexity of the data and can be 2 in one embodiment of the present application.
[0069] Figure 4 is a schematic diagram of the model training process provided in the embodiment of the present application, such as Figure 4 As shown in the figure, the commonly used minimization of Triplet Loss in the twin network is used for training to learn representations with strong separability and discriminability. Triplet Loss effectively learns a more fine-grained feature representation, captures the similarities and differences between samples, and helps to achieve a more refined distinction between similar roots.
[0070] Specifically, a Triplet consists of a difference sample (anchor point), a similar difference sample (positive sample), and a different difference sample (negative sample). The number of combinations is very large, which makes the training data grow exponentially. However, due to the twin structure, the number of parameters will not increase significantly, effectively alleviating the overfitting of the model. Triplet Loss is defined as:
[0071] in, , and Represent anchor points, positive samples and negative samples respectively, represents the Euclidean distance, is a hyperparameter, usually set to 1.0.
[0072] It is very inefficient to traverse all triplets during training. Therefore, in order to accelerate network convergence, this method adopts the Online Triplet Generation Strategy proposed in FaceNet. Specifically, in each mini-batch, for each difference sample , we randomly select a difference sample of the same type as a positive sample , and select its hard negative sample Perform gradient descent.
[0073] In some embodiments, step 130 specifically includes: Step 1301, obtaining a difference sample, where the difference sample is a pair formed by an abnormal sample and a normal sample in the database node cluster to be tested; Step 1302: Input the difference samples into the trained twin difference network to determine the abnormal root cause type of the database node cluster to be tested.
[0074] Figure 5 is a schematic diagram of the process of the diagnostic test sample provided in the embodiment of the present application, such as Figure 5 As shown, to diagnose a test sample, the goal is to identify similar samples to it in the training set, since these samples should have the same root cause as the test sample.
[0075] First, the sample construction method is applied to pair the test sample with the normal sample of the corresponding cluster, thus obtaining a set of In order to ensure the consistency of the inference results, only the clustering-based construction method is considered. The difference samples are input into the trained twin difference network to generate A difference representation.
[0076] Based on the difference representation, the closest training set sample can be determined, and then the root cause type can be determined according to the label of the sample to determine the final abnormal root cause type.
[0077] In some embodiments, step 1302 specifically includes: Step 13021, input the difference sample into the trained twin difference network to generate a difference representation corresponding to the difference sample; Step 13022: Determine the abnormal root cause type of the database node cluster to be tested based on the difference representation and the prototype difference representation of the training sample.
[0078] like Figure 5 As shown, since each abnormal sample has In order to improve efficiency and alleviate the impact of outliers, the representation with the smallest total distance to other representations is defined as the prototype difference representation of abnormal samples, expressed as , Represents a set of difference representations.
[0079] Optionally, given a test sample The prototype difference representation of the training samples and the prototype difference representation of the training samples can be used to determine the abnormal root cause type through two strategies, as shown below: 1. Prototype-based strategy: Calculate the test sample . Then select in the training set The sample closest to the test sample. If the Euclidean distance is less than the threshold , then the test sample is diagnosed as the same root cause type as the selected sample, otherwise it is diagnosed as an unknown root cause type.
[0080] 2. Election-based strategy: For each differential representation of a test sample, identify the prototype differential representation of the nearest training sample. Then perform a majority vote based on the labels of these nearest neighbors. If more than half of the neighbors have the same root cause type as the test sample, the test sample is diagnosed with that root cause type, otherwise it is an unknown root cause type. The idea behind this is that the differential representation of samples with root causes is consistent with various normal patterns. Specifically, the differential representations between abnormal samples with known / unknown root cause types and different normal samples tend to cluster / scatter in the feature space.
[0081] In some embodiments, the method further comprises: Step 1303, transform the difference sample and input it into the trained twin difference network to obtain the difference representation corresponding to the transformed sample; Step 1304, obtaining a symptom KPI set based on the difference representation corresponding to the transformed sample and the deviation of the difference representation corresponding to the difference sample; Step 13033, calculate the Personalized PageRank Vector corresponding to each KPI in the symptom KPI set, and determine the root cause KPI that causes the abnormality of the database node cluster to be tested based on the Personalized PageRank Vector.
[0082] To locate the root cause KPI, you need to locate the symptom KPI first, because the root cause KPI is usually located in the symptom KPI set.
[0083] First locate the symptom KPI, given the prototype difference sample of the test sample , for each , the abnormal samples The time series of is replaced by the time series in the normal sample, and the result is recorded as .
[0084] Will Input the trained twin difference network and get the output difference representation. Then, calculate the deviation between the transformed difference representation and the original difference representation. .
[0085] use Instead The greater the deviation, the The higher the degree of abnormality. It can be considered as the anomaly score of the KPI. The reason behind this is that for normal KPIs in abnormal samples, their patterns are similar to those of normal samples, so replacing them will not significantly affect the differential representation. On the contrary, replacing abnormal KPIs with normal KPIs will greatly affect the differential representation.
[0086] Optionally, you can select The largest several KPIs serve as symptom KPIs.
[0087] Then locate the root cause KPI. To further locate the root cause KPI, a weighted undirected dependency graph (WUDG) can be used to replace the causal graph. By completing the missing edges between nodes and adding the Fisher-Z test The reciprocal of is used as the edge weight, and the KPI dependency graph obtained in the process of constructing the hierarchical dependency graph is expanded to WUDG.
[0088] Then, the Personalized PageRank algorithm is used to replace the standard PageRank algorithm that treats each node equally to better capture the propagation of anomalies. Personalized PageRank allows setting preferences for specific nodes, making it more suitable for anomaly propagation scenarios.
[0089] In this method, Personalized PageRank Vector (PPV) is calculated as the root cause score of each node.
[0090] To calculate PPV, the transition probability between nodes is defined as ,in Representation Node and The normalized anomaly score of the KPI is then calculated as the node preference vector ,in , making KPIs with large abnormalities easier to access during random transfer. Formally, the PPV calculation formula on WUDG is:
[0091] in is the damping factor, usually set to , Representation and The set of directly connected nodes.
[0092] Finally, several symptom KPIs with the highest PPV can be used as root cause KPIs.
[0093] Figure 6 is a schematic diagram of the structure of an abnormal intelligent diagnosis system based on a distributed cloud database provided in an embodiment of the present application, such as Figure 6 As shown, the system includes a construction module 610, an acquisition module 620, a training module 630 and a determination module 640, wherein: A construction module 610 is used to construct a hierarchical dependency graph based on the key performance indicator KPI time series data of the database node cluster; An acquisition module 620 is used to acquire training samples, where the training samples include abnormal samples and samples formed by pairing normal samples in a cluster corresponding to the abnormal samples; A training module 630 is used to train the twin difference network based on the training samples and the hierarchical dependency graph using a Triplet Loss loss function to obtain a trained twin difference network; The determination module 640 is used to determine the abnormal root cause type of the database node cluster to be tested based on the trained twin difference network.
[0094] It should be understood that the above-mentioned system is used to execute the methods in the above-mentioned embodiments. The implementation principles and technical effects of the corresponding program modules in the system are similar to those described in the above-mentioned methods. The working process of the system can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0095] Based on the method in the above embodiment, Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7 As shown, an embodiment of the present application provides an electronic device, which may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the abnormal intelligent diagnosis method based on the distributed cloud database in the above embodiment.
[0096] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the abnormal intelligent diagnosis method based on a distributed cloud database described in each embodiment of the present application.
[0097] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the abnormality intelligent diagnosis method based on a distributed cloud database in the above embodiment.
[0098] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the abnormality intelligent diagnosis method based on a distributed cloud database in the above embodiment.
[0099] It is understandable that the processor in the embodiment of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0100] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0101] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0102] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0103] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. An abnormal intelligent diagnosis method based on a distributed cloud database, characterized in that: include: Build a hierarchical dependency graph based on the KPI time series data of the database node cluster; Acquire training samples, where the training samples include abnormal samples and samples formed by pairing normal samples in a cluster corresponding to the abnormal samples; Based on the training samples and the hierarchical dependency graph, the twin difference network is trained using a Triplet Loss loss function to obtain a trained twin difference network; Based on the trained twin difference network, the abnormal root cause type of the database node cluster to be tested is determined.
2. The abnormal intelligent diagnosis method based on distributed cloud database according to claim 1 is characterized in that: The key performance indicator KPI time series data of the database node cluster is used to construct a hierarchical dependency graph, including: Use the Isolation Forest algorithm to clean the normal KPI time series data of the database node cluster; Based on the cleaned normal KPI time series data and the abnormal KPI time series of the database node cluster, a KPI dependency graph is constructed; The KPI dependency graph is expanded into a hierarchical dependency graph.
3. The abnormal intelligent diagnosis method based on distributed cloud database according to claim 1 is characterized in that: Based on the training sample and the hierarchical dependency graph, the twin difference network is trained using the Triplet Loss loss function to obtain a trained twin difference network, including: Constructing a twin difference network, the twin difference network is composed of three shared weight difference networks, including a multi-scale trend convolution module and a hierarchical graph convolution module, wherein the multi-scale trend convolution module is used to extract short-term and long-term trends and horizontal features from the KPI time series data, and the hierarchical graph convolution module is used to mine the correlation of the KPI time series data based on the hierarchical dependency graph; The constructed twin difference network is trained based on the training samples to obtain a trained twin difference network.
4. The abnormal intelligent diagnosis method based on distributed cloud database according to claim 1 is characterized in that: The determining the abnormal root cause type of the database node cluster to be tested based on the trained twin difference network includes: Obtaining a difference sample, where the difference sample is a pair formed by an abnormal sample and a normal sample in the database node cluster to be tested; The difference sample is input into the trained twin difference network to determine the abnormal root cause type of the database node cluster to be tested.
5. The abnormal intelligent diagnosis method based on distributed cloud database according to claim 4 is characterized in that: The step of inputting the difference sample into the trained twin difference network to determine the abnormal root cause type of the database node cluster to be tested includes: Inputting the difference sample into the trained twin difference network to generate a difference representation corresponding to the difference sample; Based on the difference representation and the prototype difference representation of the training sample, the abnormal root cause type of the database node cluster to be tested is determined.
6. The abnormal intelligent diagnosis method based on distributed cloud database according to claim 4 is characterized in that: The method further comprises: Transforming the difference sample and inputting it into the trained twin difference network to obtain the difference representation corresponding to the transformed sample; Obtaining a symptom KPI set based on the difference representation corresponding to the transformed sample and the deviation of the difference representation corresponding to the difference sample; The Personalized PageRank Vector corresponding to each KPI in the symptom KPI set is calculated, and the root cause KPI causing the abnormality of the database node cluster to be tested is determined based on the Personalized PageRank Vector.
7. An abnormal intelligent diagnosis system based on a distributed cloud database, characterized in that: include: A construction module is used to build a hierarchical dependency graph based on the key performance indicator (KPI) time series data of the database node cluster; An acquisition module, used to acquire training samples, wherein the training samples include abnormal samples and samples formed by pairing normal samples in a cluster corresponding to the abnormal samples; A training module, used to train the twin difference network based on the training samples and the hierarchical dependency graph using a Triplet Loss loss function to obtain a trained twin difference network; A determination module is used to determine the abnormal root cause type of the database node cluster to be tested based on the trained twin difference network.
8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the abnormal intelligent diagnosis method based on the distributed cloud database as described in any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor executes the abnormal intelligent diagnosis method based on a distributed cloud database as described in any one of claims 1 to 6.
10. A computer program product, characterized in that When the computer program product runs on a processor, the processor executes the abnormal intelligent diagnosis method based on a distributed cloud database as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Supply chain exceptional event detection method based on twin neural network
CN112465045A
Large-scale point cloud scene recognition method based on discriminative region feature learning
CN118865298A
Cited By
Data processing method and device, electronic equipment, storage medium and program product
CN120470043A