Abnormality detection method and device
By building the topology diagram of the target service and analyzing the node data characteristics, selecting appropriate anomaly detection algorithms, and performing abnormal detection on the target service under the distributed architecture, solving the problem of low abnormal detection efficiency and accuracy under the distributed architecture, and achieving efficient abnormal root cause positioning.
Patent Information
- Application Number
- CN202510122346.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
Under a distributed architecture, the system complexity increases, resulting in a significant increase in the difficulty of abnormal detection, and the efficiency and accuracy of abnormal detection cannot be guaranteed.
By building a topology diagram based on the topology data of the target service, obtaining the node data of the topology node, analyzing the data characteristics and determining the corresponding anomaly detection algorithm, performing abnormality detection on the topology nodes, and determining the root cause of the abnormality of the target service.
It improves the accuracy and efficiency of abnormal detection for target services, can more effectively locate the abnormal root cause and improve the high availability of the system.
Smart Images

Figure CN119988079A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present specification relate to the field of computer technology, and more particularly to anomaly detection methods and devices. Background Art
[0002] With the rapid development of the financial industry, financial institutions such as banks, insurance companies, and securities companies are gradually optimizing the distributed architecture of their task systems in order to improve the flexibility and scalability of the systems. This transformation process poses severe challenges to anomaly detection methods. In the traditional monolithic application architecture, the system components of the monolithic application are relatively concentrated, and the functional modules and logic modules are deployed in an integrated manner, which makes the anomaly detection process relatively intuitive and simple. However, with the widespread application of distributed architecture, the complexity of the system has increased significantly. A seemingly simple fault may often involve multiple application (or service) components and system resources, which greatly increases the difficulty of anomaly detection and cannot guarantee the efficiency and accuracy of anomaly detection. Therefore, there is an urgent need for a more effective anomaly detection method to deal with the above problems. Summary of the invention
[0003] In view of this, an embodiment of this specification provides an anomaly detection method. One or more embodiments of this specification also relate to an anomaly detection device, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.
[0004] According to a first aspect of an embodiment of this specification, there is provided an anomaly detection method, comprising: Construct a topology map based on the topology data of the target service, and obtain node data of the topology nodes in the topology map; Analyzing data characteristics of the node data, and determining an anomaly detection algorithm corresponding to the topological node based on the data characteristics; The anomaly detection algorithm is used to perform anomaly detection on the topological node, and a target topological node corresponding to the abnormal root cause of the target service is determined in the topological node according to the detection result.
[0005] According to a second aspect of the embodiments of this specification, there is provided an abnormality detection device, including: A construction module is configured to construct a topology map based on the topology data of the target service and obtain node data of the topology nodes in the topology map; A determination module, configured to analyze data characteristics of the node data, and determine an anomaly detection algorithm corresponding to the topological node based on the data characteristics; The detection module is configured to perform anomaly detection on the topological nodes using the anomaly detection algorithm, and determine a target topological node corresponding to the abnormal root cause of the target service in the topological nodes according to the detection result.
[0006] According to a third aspect of an embodiment of this specification, a computing device is provided, including: Memory and processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the above-mentioned abnormality detection method are implemented.
[0007] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned abnormality detection method are implemented.
[0008] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions, which implement the steps of the above-mentioned anomaly detection method when executed by a processor.
[0009] An anomaly detection method provided by an embodiment of the present specification builds a topology map based on the topology data of the target service and obtains the node data of the topology node in the topology map. The data characteristics of the node data are analyzed, and the anomaly detection algorithm corresponding to the topology node is determined based on the data characteristics. The anomaly detection algorithm matching the data characteristics is selected according to the data characteristics of the node data, thereby improving the accuracy of anomaly detection for the target service. The anomaly detection algorithm is used to perform anomaly detection on the topology node, and the target topology node corresponding to the abnormal root cause of the target service is determined in the topology node according to the detection result, thereby improving the efficiency of anomaly detection for the target service. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a schematic diagram of a processing process of an abnormality detection method provided by an embodiment of this specification; Figure 2 is a flow chart of an anomaly detection method provided by an embodiment of this specification; Figure 3 is a processing flow chart of an abnormality detection method provided by an embodiment of this specification; Figure 4 is a schematic diagram of the root cause location principle of an abnormality detection method provided by an embodiment of this specification; Figure 5 is a schematic diagram of an anomaly detection method provided by an embodiment of this specification; Figure 6 is a structural schematic diagram of an abnormality detection device provided by an embodiment of this specification; Figure 7 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0010] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0011] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0012] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0013] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0014] First, the terms involved in one or more embodiments of this specification are explained.
[0015] Indicators: Data or variables that measure the system performance and health of the target object, such as CPU usage, memory consumption, response time, etc. These indicators can reflect key information such as the system's operating status, resource usage, and task processing efficiency. By monitoring and analyzing these indicators, operation and maintenance personnel can promptly discover system anomalies and effectively locate and handle the root causes.
[0016] Anomaly detection: Anomaly detection aims to identify observations that do not conform to the expected pattern and obtain outliers. Outliers may appear in the form of outliers, noise, or deviations. Through anomaly detection, potential system errors or logical anomalies in the target object can be discovered in a timely manner, so as to effectively locate and deal with the root cause.
[0017] Application topology: A graph structure that describes the calling and deployment relationships between applications (or target services). This type of graph structure is called a topology graph.
[0018] Root cause location: The process of determining the initial cause of a problem by analyzing fault phenomena, logs, monitoring data, and other information. The initial cause of the problem is the root cause. Root cause location refers to finding the graph node in the topology graph of the application topology that causes the fault.
[0019] Autoregressive Model: Autoregressive model is a prediction method based on time series. It uses the observed values at past time points to predict the values at future time points. It establishes the relationship between the current value and the values at previous times, thus capturing the autocorrelation of the time series.
[0020] Moving Average Model: Moving average model is a statistical analysis method. The data within a certain period is averaged and the average values at different times are connected to observe the trend of data changes. The main function of the moving average model is to smooth the data and eliminate random fluctuations in the time series, making it easier to identify long-term trends or cyclical patterns.
[0021] Support Vector Machine: Support Vector Machine is a machine learning algorithm used for classification and regression analysis. It maximizes the interval between two types of samples by finding a hyperplane. In time series signal analysis, it can be used to handle classification and regression tasks.
[0022] Gradient Boosting Machines: Gradient Boosting Machines is an ensemble learning algorithm. It corrects the error of the previous model by continuously building new decision trees. In time series signal analysis, it can be used to perform regression and classification tasks.
[0023] Recurrent Neural Network: Recurrent Neural Network is a neural network model used to process sequence data. By introducing recurrent connections, the network can capture the temporal dependencies in sequence data.
[0024] Anomaly confidence: An indicator used to quantify the likelihood that a topology node is considered to be the root cause of anomalies. The higher the anomaly confidence, the greater the likelihood that the topology node is considered to be the root cause of anomalies.
[0025] K-Sigma algorithm: Also known as the K times standard deviation algorithm, it is a method for detecting outliers in a data set. It is based on the assumption that the values of normal data are concentrated around a mean and its variance is relatively stable. In the K-Sigma algorithm, the mean (μ) and standard deviation (σ) of the data set are first calculated. Then, a K value (usually a positive integer) is selected to define the range of normal data as the interval [μ - Kσ, μ + Kσ]. If a data point is outside this interval, it is considered an outlier.
[0026] N-Sigma algorithm: An anomaly detection method based on statistical principles. Its core idea is to determine the normal range of data by setting a certain standard deviation multiple (Sigma value, i.e. σ) and identify outliers outside this range. In a normal distribution, the probability that a data point falls within a certain interval near the mean (μ) is known. Specifically, about 68.27% of the data points fall within the range of mean ±1σ, about 95.45% of the data points fall within the range of mean ±2σ, and about 99.73% of the data points fall within the range of mean ±3σ. The N-Sigma algorithm is based on this principle. By setting an N value (usually a positive integer, such as 1, 2, 3, etc.), the range of normal data is defined as the interval [μ - Nσ, μ + Nσ]. If a data point exceeds this interval, it is considered an outlier.
[0027] Prophet algorithm: an open source time series forecasting algorithm, the core idea of which is to decompose time series data into components such as trend term, seasonal term, holiday term and residual term (error term), and predict future time series values by fitting these components. Among them, the trend term (Trend) is used to simulate the non-periodic changes of time series values. The seasonal term (Seasonality) is used to capture the periodic changes in the time series. The holiday term (Holidays) considers the impact of holidays or special events on the time series. The residual term (Residual) represents special changes that the model is not adapted to, and it is usually assumed to follow a normal distribution.
[0028] XGBoost (eXtreme Gradient Boosting): An optimized distributed gradient boosting library designed for efficient, flexible, and portable machine learning models. The core idea of XGBoost is a boosting algorithm based on the Gradient Boosted Decision Tree (GBDT).
[0029] CMDB (Configuration Management Database): A centralized database used to store the configuration, attributes, and relationships of all important information technology assets (such as servers, network devices, etc.) within an information system. Important information included includes: all applications (services) in the environment, the computer room where the application is deployed, the container where the application is deployed, virtual machines, physical machines, etc.
[0030] Responsive: Responsive programming is a programming paradigm oriented towards data streams and change propagation, aiming to build dynamically updated systems by responding to changes in these data streams.
[0031] In the process of informatization of financial industries, such as banks, insurance and securities institutions, with the popularization and deepening of distributed computing and containerized deployment modes, troubleshooting methods in system operation and maintenance are facing challenges. In an environment where functional modules are tightly coupled and deployed, it is easier to locate the source of the problem when a failure occurs. However, after turning to a distributed architecture, the application is decomposed into multiple independent service components, and the services interact through the network and may be distributed on different physical or virtual servers. Such systems usually rely on the support of various middleware and other system resources. Therefore, once a failure occurs, its scope of impact may span multiple service components and service levels, making it complicated and difficult to determine the root cause of the problem. Faced with massive application data and intricate call links, there is a lack of an effective means to automatically analyze and infer the root cause of the failure, which has become a major obstacle to improving the high availability of the system. An embodiment of this specification provides an anomaly detection method to solve the above problems.
[0032] Figure 1 FIG. 1 is a schematic diagram of a processing process of an abnormality detection method provided by an embodiment of this specification; Figure 1 As shown, the target service can be a service provided by a system with complex and large calls and dependencies. During the operation of the target service, the parameter values generated during the operation of the target service can be detected in real time. When the parameter value exceeds the preset parameter value range, the target service triggers an alarm. At this time, it is necessary to locate the root cause of the abnormality of the target service based on this alarm to determine the root cause of the abnormality of the target service. When locating the root cause of the abnormality of the target service, indicator collection and link tracking can be performed for the target service, and indicator data on the target service and related applications and deployment machines can be collected, such as CPU utilization, memory usage, ports and other data. Call data such as call relationships and call time between target services and applications are collected. Topological data of the target service is constructed based on indicator data and call data. A topological graph is constructed based on the topological data of the target service, and node data of each topological node in the topological graph is obtained. Data characteristics of node data are analyzed for topological nodes, and data characteristics include but are not limited to periodic characteristics and seasonal characteristics.
[0033] The anomaly detection algorithm corresponding to the topological node is determined based on the data characteristics. When selecting the anomaly detection algorithm, different types of anomaly detection algorithms can be selected according to the different data characteristics of the node data. When the node data has periodicity, it means that the node data has a periodic trend. You can select an anomaly detection algorithm based on the periodicity of the data to realize anomaly detection of the topological nodes in the topological map by processing the node data. When the node data does not have periodicity, you can select an anomaly detection algorithm that detects outliers in the data set to perform anomaly detection on the topological nodes in the topological map. According to the data characteristics of the node data, select an anomaly detection algorithm that matches the data characteristics to improve the accuracy of anomaly detection for the target service. Use the anomaly detection algorithm to detect anomalies on the topological nodes, and determine the target topological node corresponding to the abnormal root cause of the target service in the topological node according to the detection results, thereby improving the efficiency of anomaly detection for the target service.
[0034] In this specification, an anomaly detection method is provided. This specification also relates to an anomaly detection device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0035] See also Figure 2 , Figure 2 A flowchart of an abnormality detection method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0036] Step 202: construct a topology map based on the topology data of the target service, and obtain node data of the topology nodes in the topology map.
[0037] Specifically, the target service can be a service provided by a system with complex and large calling and dependency relationships. The target service can be a financial service provided by financial fields such as banking, insurance, and securities, or a service provided by applications such as teaching, social networking, and shopping. The topological data can be the calling relationship data and deployment relationship data between the sub-services in the target service, and the topological graph is a graph structure that describes the calling relationship and deployment relationship between the sub-services in the target service. The topological node can be a graph node representing a sub-service in the topological graph, and the node data can be the sub-service data of the sub-service corresponding to the topological node. The node data is observable data, including but not limited to the indicator data and log data of the sub-service, where the indicator data can be service indicators, link indicators, basic resource indicators and other data.
[0038] In actual applications, the target service can be an application set, which contains at least two applications with call relationships or dependency relationships. The topological data of the target service obtained can be metadata topology, which is obtained by performing real-time link cleaning, deployment architecture analysis, and infrastructure dependency analysis on the topological data. Metadata is data that describes data. It provides information about the background, content, structure, source, quality, permission management, and other related information of the data set. Metadata topology refers to the association and organizational structure between metadata. It describes how different metadata entities are connected and interact with each other, as well as their position and role in the entire data ecosystem. Obtaining topological data for topological nodes means pulling indicators for topological nodes, pulling task indicators, link indicators, basic resource indicators, application service indicators, events, and logs corresponding to the topological nodes. So that when an alarm is generated by the target service in the future, the target service can be detected and the root cause of the abnormality can be located based on the node data.
[0039] In specific implementation, the complete service link of the target service can be obtained through real-time link cleaning. The order link in the target service needs to flow through at least one application. The order links generated within a period of time can be collected, and the complete order link can be obtained by order link merging and order link cleaning. For example, the order links collected within a period of time include order link 1: application A-application B; order link 2: application A-application C; order link 3: application C-application D. After order link cleaning and merging of the above order link 1, order link 2, and order link 3, and deleting the repeated call relationships in the four order links, the complete order link can be obtained: application A-application C-application D, application A-application B. This complete order link represents the call relationship of the target service.
[0040] You can also determine the dependencies of the target service by analyzing the deployment architecture corresponding to the target service. Application A, Application B, and Application C above have deployed at least one instance. Instances of the target service can be deployed on containers or hosts, and the deployment information can be stored in the CMDB. Generally, this information is automatically collected during deployment. For example, if you use k8s (Kubernetes, an open source container orchestration and management platform) for deployment, you can automatically collect how many instances of each application are and on which nodes they are deployed. The CMDB can also maintain information on infrastructure such as databases, and use this information as infrastructure dependencies for subsequent target service anomaly detection and anomaly root cause location.
[0041] Furthermore, considering that the sub-services in the target service have complex calling relationships and dependency relationships, it is necessary to fully collect the topological data of the target service when constructing a topological map. Specifically, the topological map is constructed based on the topological data of the target service, including: determining at least two target applications included in the target service, and application relationship data corresponding to the at least two target applications; determining graph nodes based on the at least two target applications, and determining graph relationships based on the application relationship data corresponding to the at least two target applications; and constructing the topological map using the graph nodes and the graph relationships as the topological data.
[0042] Among them, the target application refers to the application corresponding to the sub-service in the target service, and refers to the functional application provided in the target service. Application relationship data refers to the relationship data between one target application and other target applications in at least two target applications. The relationship between target applications can be a call relationship or a dependency relationship. Graph nodes refer to topological nodes in a topological graph, and target applications can be topological nodes in a topological graph. Graph relationships represent dependency relationships or call relationships between target nodes. Graph relationships can be represented by directed lines, and the constructed topological graph can be a directed acyclic graph. A directed acyclic graph is a data structure composed of vertices and edges, in which each edge has a clear direction, and the entire graph is acyclic, that is, there is no path in the graph that can start from a vertex, pass through a series of edges, and then return to the vertex.
[0043] In actual applications, when constructing a topology map, it is necessary to identify at least two target applications included in the target service. These applications constitute the basic unit of analysis. In order to deeply understand the interactions between these applications, it is necessary to obtain the application relationship data corresponding to them. These data usually include key information such as the call relationship between applications and the direction of data flow. Determine the graph nodes in the topology map based on these target applications. Each target application is regarded as an independent node, representing a component of the service architecture. Then use the obtained application relationship data to determine the graph relationships between nodes. These relationships reveal how applications collaborate and interact with each other. After determining the nodes and relationships, you can draw a topology map. Draw the graph nodes corresponding to at least two target applications, which are usually presented as circles in the map. Draw the directed lines between the nodes based on the previously determined relationships. Through this series of steps, you can finally build a complete topology map, which intuitively shows the overall architecture of the target service and the interaction relationship between each application.
[0044] For example, when the target service is a payment service, the target service contains three applications, A, B, and C. Application A provides an ordering interface. If the order is successful, it calls B for subsequent operations; if the order fails, it calls C to send an email notification. For a specific call, its call topology can be A->B, or A->C. These two topologies are incomplete in describing the ordering operation. By obtaining the call link and cleaning all calls from an interface, a clear horizontal topological relationship of the application can be cleaned, that is, C<-A->B. The three applications A, B, and C also have their own deployment dependencies, that is, the instance of the virtual machine (container) deployed by each application, the host that carries the virtual machine (container), the computer room, and other information. By merging the topological relationship and deployment dependency information, a topological map with clear semantics can be obtained. A complete topological map is constructed for the target service to improve the comprehensiveness of anomaly detection for the target service.
[0045] Furthermore, during the operation of the target service, log data will be generated in real time, and the indicators of each service will also change continuously. When obtaining the node data of a topological node, the indicator data and log data of the topological node can be obtained, and the indicator data and log data are used as the node data. Specifically, obtaining the node data of the topological node in the topological map includes: determining at least two topological nodes in the topological map; and using the indicator data and log data corresponding to the at least two topological nodes as the node data.
[0046] Among them, the topological nodes in the topological map represent the target applications corresponding to the sub-services under the target service. The node data can be obtained for all the topological nodes in the topological map. For the topological nodes in the topological map, the data corresponding to each topological node is obtained to form the node data. The indicator data is used to measure the performance and health of the system, which can be data such as CPU usage, memory consumption, response time, etc. Log data refers to the record information automatically generated during the operation of the target service, which records key information such as various events, state changes, and user activities.
[0047] In practical applications, the topology map contains at least two topology nodes. Data can be collected for all topology nodes contained in the topology map, and the indicator data and log data corresponding to at least two topology nodes are collected. The indicator data and log data corresponding to at least two topology nodes are used as the node data collected for the target service. Subsequently, the root cause of abnormality of the target service can be detected and located based on the node data.
[0048] The topology map can be used as a visual representation of the architecture of the target service. A complete topology map contains at least two topology nodes, which represent different components or applications in the architecture of the target service. In order to deeply understand the status and behavior of these components or applications, it is necessary to collect data from all topology nodes in the topology map. Data collection is a basic step in detecting and locating the root cause of anomalies. For each topology node, its corresponding indicator data and log data will be collected. Indicator data usually includes key performance indicators such as CPU usage, memory usage, and disk I / O, which can reflect the operating status and health of the node. The log data records various events and information generated by the node during operation, including error logs, warning logs, etc., which play a vital role in locating the root cause of the problem.
[0049] Integrate the collected indicator data and log data to form node data for the target service. These data provide a rich source of information, based on which we can have a more comprehensive understanding of the overall status of the target service and the behavior of each component. Based on the node data, detect and locate the root cause of abnormalities in the target service. By analyzing the abnormal patterns in the node data, the changing trends of the performance indicators, and the key information in the logs, we can gradually narrow the scope of the problem and eventually locate the root cause of the service abnormality. This process not only improves the efficiency of problem solving, but also provides valuable experience for optimizing the service architecture and improving operation and maintenance practices. By collecting data for all topological nodes in the topology diagram and integrating them to form node data, a solid foundation can be laid for subsequent abnormal root cause detection and location work.
[0050] Using the above example, the target service includes three applications, A, B, and C. The indicator data and log data can be collected for A, B, and C respectively. The collected indicator data include service indicators, link indicators, basic resource indicators, etc. The indicator data and log data collected for A, B, and C are used as node data for subsequent anomaly detection. Fully collect node data for the topology map to improve the accuracy of subsequent anomaly detection for the target service.
[0051] Step 204: Analyze the data characteristics of the node data, and determine the anomaly detection algorithm corresponding to the topological node based on the data characteristics.
[0052] Specifically, after constructing a topology map based on the topology data of the target service and obtaining the node data of the topology nodes in the topology map, the data characteristics of the node data can be analyzed, and the anomaly detection algorithm corresponding to the topology node can be determined based on the data characteristics, wherein the data characteristics of the node data represent the data attributes of the node data in the data distribution dimension, and the data characteristics include but are not limited to data periodicity and data seasonality, which can reflect the regularity of the node data. The anomaly detection algorithm is used to perform anomaly detection on the topology nodes in the topology map and detect the root cause topology node that causes the anomaly.
[0053] Based on this, after constructing a topological map based on the topological data of the target service and obtaining the node data of the topological nodes in the topological map, the data characteristics of the node data in the data distribution dimension are analyzed, and an anomaly detection algorithm with a higher degree of match with the topological node is determined based on the data characteristics, so that the anomaly detection algorithm can be used to locate the root cause of the anomaly in the topological node in the topological map in the subsequent process.
[0054] In practical applications, node data often show obvious periodic or seasonal characteristics, which are derived from changes in various natural and socio-economic factors. In view of this characteristic of node data, it is particularly important to select a suitable anomaly detection algorithm. In order to more effectively identify outliers in node data, an anomaly detection algorithm with a high degree of match with the data characteristics of the node data can be selected. This type of algorithm can make full use of the periodicity or seasonality of node data, and accurately capture abnormal changes by comparing the difference between current data and historical data. Doing so can not only improve the accuracy of anomaly detection, but also reduce false positives and false negatives, providing a more reliable basis for subsequent decision analysis. In actual operation, combined with the specific characteristics of node data, the anomaly detection algorithm can be flexibly selected or optimized to ensure the effectiveness and efficiency of anomaly detection.
[0055] Furthermore, timing signal analysis can be performed on the node data to determine data characteristics of the node data. Specifically, the analysis of the data characteristics of the node data includes: performing timing signal analysis on the node data; when it is determined according to the analysis result that the node data contains missing data, determining the data characteristic to be data missing; when it is determined according to the analysis result that the node data does not contain missing data, performing signal period processing on the node data, and determining the data characteristics of the node data based on the processing result.
[0056] Among them, the purpose of performing time series signal analysis on node data is to analyze whether the node data contains missing data. Missing data refers to a continuous period of data in the node data that contains some null value data. The reason for the missing data in the node data may be technical failure, such as: data acquisition equipment failure, data transmission problem, data storage problem and system operation error, or natural factors (temperature and humidity affect the operation of the equipment) leading to data missing. Statistical methods, such as autoregressive model, moving average model, etc., can be used to perform time series signal analysis on node data. Machine learning algorithms and deep learning models can also be used to perform time series signal analysis on node data. Machine learning algorithms include but are not limited to support vector machines and gradient boosting machines; deep learning modules include but are not limited to recurrent neural networks.
[0057] In an autoregressive model, the current value is represented as a linear combination of several past values plus a random error term. The key to the model is to determine the appropriate number of lag terms and the corresponding coefficients, which is usually achieved by fitting the data and optimizing the model parameters. Once the model is established, past time series data can be used to predict future values and realize time series signal analysis of node data. The main function of the moving average model is to smooth the data and eliminate random fluctuations in the time series, making it easier to identify long-term trends or periodic patterns. It is suitable for time series data that is greatly affected by periodic and irregular changes, and can realize time series signal analysis of node data.
[0058] Support vector machine (SVM) is a classification and regression algorithm based on statistical learning theory. Its core idea is to maximize the interval between different categories by finding an optimal hyperplane, thereby improving the generalization ability of the classifier. In time series signal analysis, SVM can be used for time series signal analysis through classification and regression. SVM can be used for classification tasks (such as identifying different patterns or events in the signal) and regression tasks (such as predicting the future value of the signal). Gradient boosting machine (GBM) is an ensemble learning algorithm based on decision trees. Its core idea is to gradually optimize the model in an iterative manner to improve the prediction accuracy. Recurrent neural network (RNN) is a neural network specifically used to process sequence data. Its core lies in the recurrent connection, which enables the network to keep the memory of previous information and capture the dynamic features in the sequence. The output of RNN will be used as the input of the next time step to form a recurrent connection, so that the network can remember the previous input information. This feature makes RNN perform well in processing time series data. In addition, RNN can remember the information in the sequence, which is crucial for capturing long-term dependencies in time series signals.
[0059] Signal period processing for node data can be signal period detection and period signal noise reduction. Signal period detection is used to detect whether the node data is periodic by autocorrelation analysis, while signal period noise reduction is used to perform period noise reduction on node data in combination with the characteristics of indicators. For example, the daily period should be a multiple of 1440+10% or 1440-10%. Periods that are not in this range will be eliminated.
[0060] In practical applications, the node data can be analyzed for time series signals to analyze whether the node data contains missing data. If the node data is determined to contain missing data according to the analysis results, it means that the collected node data is incomplete and contains missing values. At this time, the data characteristic is determined to be missing data. If the node data is determined not to contain missing data according to the analysis results, it means that the collected node data is relatively complete and does not contain missing values. At this time, the node data can be further processed by signal cycles, and the data characteristics of the node data can be located again according to the processing results. The data characteristics of the node data can be located according to the different data characteristics of the node data, so that the appropriate anomaly detection algorithm can be selected according to the data characteristics in the future to improve the accuracy of anomaly detection.
[0061] Furthermore, considering that whether the node data has periodicity is a relatively important feature, when the node data corresponds to a periodic signal, it indicates that the node data has periodicity, and the node data can be processed according to the periodicity of the node data. Specifically, the node data is subjected to signal periodicity processing, and the data characteristics of the node data are determined according to the processing result, including: performing signal periodicity detection on the node data, and when it is determined that the node data corresponds to a periodic signal, performing periodic signal noise reduction on the node data to obtain target node data; when the target node data corresponds to a periodic signal, determining that the data characteristics of the node data are periodic signals; and when the target node data does not correspond to a periodic signal, determining that the data characteristics of the node data are non-periodic signals.
[0062] Among them, the signal period detection of the node data is to perform autocorrelation analysis on the data indicators corresponding to the node data to determine the data indicator period of the node data. Autocorrelation analysis is a statistical method used to study the self-correlation of the node data, which is a time series data, at different time points. Simply put, it measures the correlation of the same variable at different time lags. Periodic signal denoising of the node data refers to periodic denoising combined with the characteristics of the data indicators to eliminate the periods outside the preset interval. For example, the period at the day level should be a multiple of 1440+10% or 1440-10%. If the period is not in this interval, it will be eliminated. The target node data is the node data obtained after periodic signal denoising of the node data. In the case where the target node data corresponds to a periodic signal, it means that the node data still has periodicity after periodic signal denoising, so the data characteristics of the node data are periodic signals. In the case where the target node data does not correspond to a periodic signal, it means that the node data has periodicity, but the node data loses its periodicity after periodic signal denoising. At this time, the target node data that does not have periodicity does not have periodicity, which means that the data characteristics of the node data are non-periodic signals.
[0063] In practical applications, after acquiring the node data, the node data is subjected to a signal period detection to detect whether the data indicators of the node data have periodicity. When it is determined that the node data corresponds to a periodic signal, it indicates that the data indicators of the node data have a signal period and are distributed periodically. In order to eliminate abnormal periods in the node data, the node data can be subjected to periodic signal denoising to eliminate periods that are not in the preset interval to obtain the target node data. At this time, the target node data is subjected to a signal period detection again. When it is detected that the target node data corresponds to a periodic signal, it indicates that the target node data still has periodicity, and at this time, the data characteristics of the node data are determined to be periodic signals; when the target node data does not correspond to a periodic signal, it indicates that after the node data is subjected to periodic signal denoising, the target node data obtained loses periodicity, and at this time, the data characteristics of the node data are determined to be non-periodic signals. Subsequently, different anomaly detection algorithms can be selected for different data characteristics of the node data, thereby improving the accuracy of anomaly detection for the target service.
[0064] Furthermore, considering that different node data have different data characteristics, when selecting anomaly detection algorithm for detecting anomalies on the target service, the data characteristics of the node data can be fully considered. Specifically, determining the anomaly detection algorithm corresponding to the topological node based on the data characteristics includes: when the data characteristic is data missing or a periodic signal, determining the feature detection algorithm corresponding to the topological node, and using the feature detection algorithm as the anomaly detection algorithm; when the data characteristic is a non-periodic signal, determining the baseline detection algorithm corresponding to the topological node, and using the baseline detection algorithm as the anomaly detection algorithm.
[0065] Among them, when the data characteristics are missing data or periodic signals, it means that there is missing data in the node data, or the data indicators corresponding to the node data do not have periodicity. The feature detection algorithm can use an anomaly detection algorithm based on statistics, such as the K-Sigma algorithm and the N-Sigma algorithm. The difference between the K-Sigma algorithm and the N-Sigma algorithm is that the value of n can be adjusted according to the specific situation. The value of n in the N-Sigma algorithm can be adjusted according to the characteristics of the data and the needs of anomaly detection. By adjusting the value of n, the sensitivity of anomaly detection can be controlled more flexibly. The K-Sigma algorithm is a statistical anomaly detection method that assumes that the data follows a normal distribution. In a normal distribution, most data points (usually 99.73%) fall within the range of three times the standard deviation of the mean (3-Sigma). Therefore, if a data point is outside this range, it is considered abnormal.
[0066] Feature detection algorithms can also use prediction algorithms based on time series decomposition and machine learning, such as the Prophet algorithm; or they can use decision tree algorithms based on gradient boosting, such as the XGBoost algorithm. Among them, the Prophet algorithm decomposes the time series into components such as trend terms, seasonal terms, holiday terms, and residual terms, and predicts future time series values by fitting these components. Then, it can detect outliers based on the difference between the predicted value and the actual value. XGBoost gradually approximates the target function by constructing multiple decision trees, thereby achieving accurate classification or regression of data. In anomaly detection, XGBoost can be used for binary classification tasks, that is, dividing data into normal and abnormal categories.
[0067] When the data characteristics are non-periodic signals, the data indicators corresponding to the node data are periodic. The baseline detection algorithm is used to process data indicators with periodicity. The baseline detection algorithm is a method based on statistical principles. It mainly fits a baseline through seasonal analysis and non-parametric regression methods. Data that deviate far from the baseline are considered abnormal. The algorithm has undergone long-term engineering practice in various application scenarios, continuously iterated and optimized the algorithm logic, and has good generalization. The core of the baseline detection algorithm lies in two aspects, namely seasonal analysis and non-parametric regression. First, seasonal analysis is used to detect whether the time series has periodicity. If it does, the period size is identified and the periodic term and trend term are extracted and superimposed as the baseline. If it does not exist, the trend term is directly extracted as the baseline. When extracting the trend term, the local weighted regression algorithm is used for smoothing. Finally, the upper and lower limits are set according to the residual analysis of the original data and the baseline. Data exceeding the upper and lower limits are considered abnormal. Baseline detection algorithm.
[0068] The generation process of the baseline detection algorithm includes: period detection, period identification and period verification of the data; providing a smoothing interval for data change point detection; extracting the baseline through local weighted regression smoothing; moving the upper and lower limits and correcting the baseline and the upper and lower limits to obtain the baseline detection algorithm. In the period detection stage, the baseline detection algorithm performs periodic analysis on the input time series data. This usually involves using statistical methods (such as autocorrelation functions, periodograms, etc.) to detect periodic components in the data. The purpose is to determine whether there is a significant periodic pattern in the time series. The period identification stage is used to further identify the specific size of the period when periodicity is detected. This is usually achieved by analyzing the frequency components of the periodic signal. During the identification process, the baseline detection algorithm attempts to find a period length that is more consistent with the data characteristics. The identified period needs to be verified to ensure its accuracy and stability. Verification may include checking the stability of the period (i.e., whether the period length changes over time) and the significance of the periodic signal (i.e., whether the periodic signal is strong enough to distinguish it from noise).
[0069] Change point detection refers to the time point at which the data characteristics in the time series change significantly. In the baseline detection algorithm, change point detection is used to identify sudden changes or abnormal changes in the data, which may affect the fitting of the baseline. Based on the results of change point detection, the baseline detection algorithm determines which intervals of data are stationary, that is, there are no significant change points. These stationary intervals will be used in the subsequent local weighted regression smoothing to ensure the accuracy of the baseline. In the local weighted regression smoothing stage, the baseline detection algorithm uses local weighted regression to smooth the data in the stationary interval. Local weighted regression is a non-parametric regression method that fits a regression curve based on the local data around each data point. This method is well adapted to nonlinear trends and local changes in the data. After the local weighted regression smoothing process, the baseline detection algorithm extracts a smooth baseline. This baseline represents a combination of the trend term and the periodic term of the time series data. If the time series is not periodic, the baseline only represents the trend term.
[0070] After extracting the baseline, the baseline detection algorithm calculates the residuals between the original data and the baseline. These residuals reflect the degree of deviation between the data points and the baseline. Based on the results of the residual analysis, a reasonable upper and lower limit range is determined. This range is usually set based on the distribution characteristics of the residuals (such as mean, standard deviation, etc.). The upper and lower limits are dynamically adjusted as the baseline moves to ensure that they can accurately capture abnormal data points. After initially extracting the baseline and upper and lower limits, the baseline detection algorithm may modify them according to some additional rules or conditions. For example, if the baseline has unreasonable fluctuations or jumps in certain areas, the baseline detection algorithm may smooth or refit these areas. Similarly, the upper and lower limits may also need to be adjusted according to actual conditions. For example, if the upper and lower limits are too loose or too strict, the baseline detection algorithm may fine-tune them according to the distribution characteristics of the data or the specific needs of the user.
[0071] The baseline detection algorithm can set the sensitivity of anomaly detection: high, medium, and low, corresponding to different scale values. High sensitivity means that it is very sensitive to abnormal data changes, and may detect more anomalies, but may also cause false detections; low sensitivity means that it can tolerate some data fluctuation anomalies caused by noise, and may detect fewer anomalies, but may also cause missed detections. The default scale value is equal to 4, indicating medium sensitivity.
[0072] In practical applications, when the data characteristics are missing data or periodic signals, the corresponding feature detection algorithm of the topological node is determined. The feature detection algorithm can be used as an anomaly detection algorithm to detect anomalies in node data. After performing box plot denoising and kurtosis analysis on the node data, the feature detection algorithm is used for anomaly detection. Among them, box plot denoising is used to segmentally calculate the box plot of the data indicators of the node data, and the points outside the upper and lower quartiles are selected as noise and removed. This step is to remove unreasonable extreme points, otherwise the kurtosis cannot be correctly analyzed. Kurtosis analysis is used to determine whether the data indicators of the node data are concentrated in a range. For example, most of the interface time consumption indicators will fall within a small range, such as 100-500ms, and a few will fall in other ranges. The parameter n of the N-Sigma algorithm for this indicator will be smaller, trying to fit this relatively small range. For other indicators, such as CPU utilization, the fluctuation is relatively large and there is no fixed range. The parameter n of the N-Sigma algorithm will be larger, trying to cover the fluctuation and not generate too many false positives.
[0073] Furthermore, since node data has different data characteristics, different types of anomalies that may be present in the node data can be determined based on different data characteristics. In order to improve the accuracy of anomaly detection, an anomaly detection algorithm that matches the anomaly type can be selected to perform accurate anomaly detection on the target service. Specifically, determining the anomaly detection algorithm corresponding to the topological node based on the data characteristics includes: determining the anomaly type corresponding to the topological node based on the data characteristics, and selecting the anomaly detection algorithm based on the anomaly type.
[0074] Among them, the anomaly type indicates the type of anomaly of the topological node, which can be identified by statistics or artificial intelligence algorithms. Anomaly types include but are not limited to shock, step, slow change, zero drop, frequency anomaly and trend anomaly. Shock refers to the situation where the data indicator of the target service suddenly and drastically deviates from the normal range and then quickly returns to normal. This anomaly is usually a one-time event caused by external factors or internal short-term errors, such as a sudden influx of large traffic, instantaneous hardware failure, etc. Step anomaly means that an indicator value has changed significantly in a short period of time and remains at a new level. This may be caused by configuration changes, new version deployment, or other factors that affect system behavior for a long time. Slow change anomaly refers to the process in which the data indicator of the target service gradually changes instead of immediately changing significantly. Such anomalies may not be easy to detect because they develop slowly, but may have a significant impact on the stability and performance of the system over time. Zero drop anomaly refers to a situation in which an indicator suddenly drops to zero or close to zero and then returns to normal. This situation usually indicates a short service interruption or resource unavailability problem, such as network disconnection, temporary service failure, etc. Frequency anomaly involves the frequency of events not meeting expectations, that is, too low or too high. Trend anomaly refers to an indicator in the system showing a trend that is different from historical data.
[0075] In practical applications, it is necessary to deeply understand the characteristics of data and select appropriate anomaly detection algorithms, because different data characteristics often correspond to different anomaly types. Data characteristics can include data distribution, change trend, periodicity and other aspects. For example, some data may show obvious normal distribution characteristics, while other data may have significant periodic fluctuations. These characteristics provide clues to determine the type of anomaly. After determining the type of anomaly, you can choose an anomaly detection algorithm that matches it better based on these types. There are many types of anomaly detection algorithms, and each algorithm has its own unique advantages and applicable scenarios. For example, for data that shows normal distribution characteristics, you can choose an anomaly detection algorithm based on statistics; for data with periodic fluctuations, you can consider using an algorithm based on time series analysis. Choosing an appropriate anomaly detection algorithm is crucial to obtaining accurate anomaly detection results. An anomaly detection algorithm that highly matches the anomaly type can more effectively identify potential anomalies and reduce the possibility of false positives and false negatives.
[0076] Step 206: Perform anomaly detection on the topological nodes using the anomaly detection algorithm, and determine a target topological node corresponding to the anomaly root cause of the target service in the topological nodes according to the detection result.
[0077] Specifically, after analyzing the data characteristics of the node data as described above and determining the anomaly detection algorithm corresponding to the topological node based on the data characteristics, the anomaly detection algorithm can be used to perform anomaly detection on the topological node, and the target topological node corresponding to the abnormal root cause of the target service can be determined in the topological node according to the detection result, wherein, when performing anomaly detection on the topological node, the type of anomaly can be detected, anomaly detection is performed on each topological node in the topological graph, and anomaly confidence is assigned to each topological node, and the anomaly type and anomaly confidence are used as the detection result. A topological node with a higher confidence can be selected as the target topological node representing the abnormal root cause of the target service. The target topological node represents the topological node that originally caused the anomaly.
[0078] In practical applications, anomaly detection of topological nodes is a complex and critical task. Advanced anomaly detection algorithms can be used to accurately identify anomaly types based on statistical principles or artificial intelligence algorithms. These algorithms can automatically discover abnormal patterns or deviations in data through in-depth analysis of a large amount of node data, thereby quickly locating potential problems. Once the anomaly type is determined, the next key step is to further identify the root cause of the anomaly of the target service under that anomaly type. This usually requires in-depth analysis of the service architecture and understanding the correlation and dependency between various topological nodes. By comprehensively analyzing the node data, the scope can be gradually narrowed down, and finally the target topological node that causes the anomaly can be locked.
[0079] In specific implementation, when using artificial intelligence algorithms to detect anomalies on topological nodes, after selecting an algorithm for anomaly detection on topological nodes, an anomaly detection model is constructed. By training the anomaly detection model, the anomaly detection model completed by real-time training has the ability to detect anomalies on topological nodes. When training the anomaly detection algorithm, an anomaly detection model with good anomaly detection capabilities can be obtained through three stages of training, namely model training, parameter tuning, and model evaluation. In the model training stage, the model is trained using the selected algorithm and preprocessed data. During the training process, attention should be paid to issues such as model convergence and overfitting. After obtaining the trained anomaly detection model, the hyperparameters of the trained anomaly detection model are tuned by methods such as cross-validation, grid search, or random search. This helps to improve the generalization ability and accuracy of the model. Further, the model is evaluated, that is, the performance of the model is evaluated using indicators such as accuracy and recall.
[0080] Furthermore, when performing anomaly detection on topological nodes in a topological map, considering that the topological map includes at least two topological nodes, it is necessary to perform anomaly detection on each topological node, and use confidence to represent the possibility that each topological node is the root cause of the anomaly. Specifically, performing anomaly detection on the topological nodes using the anomaly detection algorithm includes: performing anomaly detection on the topological nodes using the anomaly detection algorithm to determine the anomaly confidence of the topological nodes; and using the anomaly confidence as the detection result.
[0081] Among them, the anomaly confidence of a topological node indicates the possibility that the topological node is the root cause of the anomaly. The anomaly confidence is an indicator used to quantify the possibility that the topological node is considered to be the root cause of the anomaly. The higher the anomaly confidence, the greater the possibility that the topological node is considered to be the root cause of the anomaly. The anomaly confidence is a value obtained by normalizing and weighting multiple dimensional data such as link depth, indicator similarity, and anomaly duration corresponding to the topological node.
[0082] In practical applications, an anomaly detection algorithm is used to perform anomaly detection on each topological node in the topological graph to determine the anomaly confidence of each topological node. The calculated anomaly confidence of each topological node is used as the detection result. The anomaly confidence of a topological node can represent the possibility that the topological node is the anomaly root cause of the target service. Based on the anomaly confidence of each topological node, the target topological node representing the anomaly root cause of the target service can be determined in the topological graph.
[0083] Using the above example, when anomaly detection is performed on three applications A, B, and C and it is determined that all three applications A, B, and C have anomalies, the anomaly confidence of each application can be further calculated. Combining the link depth, indicator similarity, and anomaly duration of the topological nodes corresponding to A, B, and C, the anomaly confidence of A is 0.3, the anomaly confidence of B is 0.6, and the anomaly confidence of C is 0.8. By combining multiple parameters to calculate the anomaly confidence of a topological node, the availability and accuracy of the anomaly confidence can be improved.
[0084] Furthermore, the target topological node corresponding to the abnormal root cause of the target service can be determined according to the abnormality confidence of each topological node in the topological map. Specifically, when there are at least two topological nodes, determining the target topological node corresponding to the abnormal root cause of the target service in the topological node according to the detection result includes: determining the abnormal topological node corresponding to the target confidence in the at least two topological nodes according to the detection result; and using the abnormal topological node as the target topological node corresponding to the abnormal root cause of the target service.
[0085] Each topological node in the topological map corresponds to an abnormal confidence. The target confidence can be a higher abnormal confidence corresponding to the topological node in the topological map, or can be two or more higher abnormal confidences corresponding to the topological node in the topological map. The topological node corresponding to the target confidence is the abnormal topological node.
[0086] In practical applications, an abnormal topological node corresponding to a target confidence is determined in at least two topological nodes according to the detection results, and the abnormal topological node is used as the target topological node of the abnormal root cause of the corresponding target service. An abnormal confidence with a higher abnormal confidence among at least two topological nodes can be used as the target confidence, and the topological node corresponding to the target confidence can be used as the target topological node. In the case where there are two or more higher abnormal confidences, both of the two or more abnormal confidences can be used as the target confidence. The topological nodes corresponding to the two or more abnormal confidences are used as the target topological nodes.
[0087] Using the above example, when the calculated abnormal confidence of A is 0.3, the abnormal confidence of B is 0.6, and the abnormal confidence of C is 0.8, the higher abnormal confidence of 0.8 can be selected. Therefore, the topological node of application C corresponding to the abnormal confidence of 0.8 can be used as the target topological node. By comparing the abnormal confidence, the target confidence is selected to improve the accuracy of abnormal root cause location.
[0088] An anomaly detection method provided by an embodiment of the present specification builds a topology map based on the topology data of the target service and obtains the node data of the topology node in the topology map. The data characteristics of the node data are analyzed, and the anomaly detection algorithm corresponding to the topology node is determined based on the data characteristics. The anomaly detection algorithm matching the data characteristics is selected according to the data characteristics of the node data, thereby improving the accuracy of anomaly detection for the target service. The anomaly detection algorithm is used to perform anomaly detection on the topology node, and the target topology node corresponding to the abnormal root cause of the target service is determined in the topology node according to the detection result, thereby improving the efficiency of anomaly detection for the target service.
[0089] The following combination Figure 3 , taking the application of the anomaly detection method provided in this specification in financial service anomaly detection as an example, the anomaly detection method is further described. Among them, Figure 3 A processing flow chart of an abnormality detection method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0090] Step 302: Obtain the call relationship and dependency relationship corresponding to the financial service.
[0091] The complexity of financial services makes the calling relationships and dependencies in financial services complex. The difficulty of financial service operation and maintenance work gradually increases with the improvement of financial services. The failure of an application or service may trigger a chain reaction, causing multiple applications or services to be abnormal. At this time, it is necessary to locate the root cause of the abnormality through a large number of indicator analyses and handle the abnormality in a timely manner.
[0092] In practical applications, the anomaly detection process in financial services is as follows: Figure 4 As shown. When an alarm is generated in a financial service, the call relationship and dependency relationship of the financial service can be obtained through link tracking and CMDB. The call relationship refers to the call relationship between applications in the application dimension of the financial service. Application A-Application B means that application A calls application B. The dependency relationship represents the deployment relationship, that is, which instances of the application are deployed and on which nodes. For example, application A has 2 instances (pod-0 and pod-1). Dependencies in financial services can correspond to the container layer and the host layer. Among them, a "container" is a lightweight, executable software package that contains the code, system tools, system libraries, and settings required to run an application. A "host" refers to a physical or virtual computer that runs a containerized application or service.
[0093] In specific implementation, the complete service link of financial services can be obtained through real-time link merging and cleaning. The order link in financial services needs to flow through at least one application. The order links generated within a period of time can be collected, and the complete order link can be obtained through order link merging and order link cleaning. For example, Figure 5 As shown in (a), the order links collected within a period of time include order link 1: application A-application B-application D; order link 2: application A-application C-application D; order link 3: application A-application C-application D and application A-application B; order link 4: application A-application C-application E. After cleaning and merging the order links 1, 2, 3 and 4, and deleting the repeated call relationships in the four order links, a complete order link can be obtained: application A-application C-application D, application A-application B, application B-application D and application C-application D. This complete order link is the complete service link of financial services in this example.
[0094] The dependencies of financial services can also be determined by analyzing the deployment architecture corresponding to financial services. The above applications A, B, C, D, and E all deploy at least one instance. Instances of financial services can be deployed on containers or hosts, and deployment information can be stored in CMDB. Generally, this information is automatically collected during deployment, such as when k8s (Kubernetes, an open source container orchestration and management platform) is used for deployment. k8s uses its core resource objects such as Deployment and ReplicaSet to automatically collect and manage the number of application instances and deployment nodes, that is, automatically collect information such as how many instances of each application there are and on which nodes they are deployed. Among them, Deployment is used to update applications and services. ReplicaSet ensures that a specified number of instances are running at any time. CMDB can also maintain information on infrastructure such as databases, and use this information as infrastructure dependencies for subsequent financial service anomaly detection and anomaly root cause location. Call relationships and dependencies can be automatically and in real time without the need for personnel to update.
[0095] Step 304: construct a topology map based on the call relationship and the dependency relationship, and analyze the financial service nodes based on the topology map to determine abnormal nodes.
[0096] After obtaining the call relationship and dependency relationship of the financial service and building the topology map, you can obtain indicator data for the financial service, and then analyze the abnormal nodes of the financial service based on the obtained indicator data.
[0097] In actual applications, the observable data middle platform can be provided as a development and deployment platform for financial services. The platform can provide a variety of service contents such as user authority management services, storage services and network services for the operation of financial services. The observable data middle platform can include monitoring, log, event, link and other modules. Among them, the link module is used to realize real-time link cleaning; the log module is used to realize log records during the operation of financial services, and the operation logs of financial services can be obtained in the log module; the event module is used to realize event records during the operation of financial services, and the events generated during the operation of financial services can be obtained in the event module; the monitoring module is used for host monitoring, container monitoring, application monitoring and network monitoring, and the task indicators, link indicators, basic resource indicators and application service indicators of financial services can be obtained in the monitoring module. Combine the indicator data and the topology map of financial services to build a topology map containing application A, application B, application C, and indicators 1-indicator 3 corresponding to each application. For each node in the topology diagram, the indicators of the application and the deployment machine (CPU utilization, memory utilization, port detection, garbage collection, etc.) are pulled, as well as the call relationship and call indicators between applications in link tracking (such as the call volume and duration of http calls), and intelligent anomaly detection is performed on each indicator.
[0098] When performing intelligent anomaly detection on each indicator, you can first intelligently identify the anomaly type of each indicator based on statistical or artificial intelligence algorithms. Anomaly types include but are not limited to shocks, steps, slow changes, zero drops, frequency anomalies, and trend anomalies. Use algorithm routing to route to different anomaly detection algorithms for anomaly detection according to the different anomaly types of indicator types. Indicators with obvious daily cycles use the baseline detection algorithm; historical data is severely missing, and the cycle and trend cannot be extracted. Use the N-Sigma algorithm; the N-Sigma algorithm is used when the indicator periodicity is very weak or there is no period; in addition, the entire root cause location process also supports manually setting a specific anomaly detection algorithm for each indicator.
[0099] The choice of anomaly detection algorithm is as follows Figure 5 As shown in (b) in the figure. For each indicator in Application A, Application B, and Application C, taking the indicator 1 of Application A as an example, we can first perform a time series signal analysis on the indicator 1 of Application A, and determine whether the indicator 1 of Application A corresponds to the problem of missing historical data based on the result of the time series signal analysis. If it is determined that the indicator 1 of Application A corresponds to the problem of missing historical data, we can perform box plot denoising on the indicator 1 of Application A, perform kurtosis analysis, and finally use the N-Sigma algorithm to perform anomaly detection on the indicator 1 of Application A.
[0100] If it is determined that there is no missing historical data for indicator 1 of application A, the signal period detection can be performed on indicator 1 of application A. If it is detected that indicator 1 of application A contains a signal period, periodic signal noise reduction can be performed. After periodic signal noise reduction, it is determined whether indicator 1 of application A is still a periodic signal. If not, the N-Sigma algorithm can still be used to detect anomalies for indicator 1 of application A. If so, it is determined that indicator 1 of application A has periodicity, and the baseline detection algorithm can be used to detect anomalies for indicator 1 of application A. The core of the baseline detection algorithm lies in two aspects, namely seasonal analysis and non-parametric regression. First, seasonal analysis is used to detect whether the time series has periodicity. If it does, the period size is identified and the periodic term and trend term are extracted and superimposed as the baseline. If it does not exist, the trend term is directly extracted as the baseline. When extracting the trend term, the local weighted regression algorithm is used for smoothing. Finally, the upper moving baseline is used as the upper and lower limits based on the residual analysis of the original data and the baseline. Data exceeding the upper and lower limits are considered abnormal.
[0101] Step 306: locate the abnormal root cause of the financial service based on the abnormal analysis result.
[0102] The above anomaly detection method is used to perform anomaly detection on indicators 1, 2 and 3 corresponding to application A, application B and application C respectively. After detection, it is determined that indicator 3 of application A, indicator 2 and indicator 3 of application C are abnormal indicators. For abnormal indicators, the root cause of abnormality can be located in multiple dimensions such as link depth, indicator similarity, and abnormal duration. The health model is used to normalize and weight the values of multiple dimensions, and finally a weighted health value is obtained. The health value can be used as the confidence of anomaly detection. The confidence values are sorted to obtain the indicator ranking of the abnormal root cause. Indicator 3 of application C is 90 points, indicator 2 of application C is 80 points, and indicator 3 of application A is 60 points. Through confidence comparison, it is determined that the abnormal root cause of financial services lies in application C, specifically, the financial service anomaly caused by indicators 2 and 3 of application C.
[0103] In the process of locating the root cause of abnormalities in financial services, the indicator pulling, abnormality detection and root cause calculation defined in the code constitute a responsive calculation graph. The root cause location is triggered by the abnormal alarm of the financial service, and all external IO calls are initiated at one time, which has a high efficiency in data pulling and calculation.
[0104] In summary, the anomaly detection method provided by an embodiment of this specification has outstanding performance in root cause location, compatibility, and data processing efficiency. Link information and CMDB information are effectively utilized, which play a vital role in the process of locating the root cause of anomalies. By integrating and deeply analyzing these diversified data, the root causes of anomalies in financial services can be more accurately identified, thereby improving the efficiency and accuracy of anomaly handling. With the help of responsive computational graph technology, high efficiency is achieved in data pulling and calculation. It ensures that a large amount of data can be processed and analyzed in a short time, thereby quickly responding to service anomalies, reducing fault recovery time, and improving the stability and reliability of the overall service.
[0105] Corresponding to the above method embodiment, this specification also provides an abnormality detection device embodiment, Figure 6 FIG. 2 shows a schematic diagram of the structure of an abnormality detection device provided by an embodiment of the present specification. Figure 6 As shown, the device comprises: A construction module 602 is configured to construct a topology map based on the topology data of the target service and obtain node data of the topology nodes in the topology map; A determination module 604 is configured to analyze data characteristics of the node data and determine an anomaly detection algorithm corresponding to the topological node based on the data characteristics; The detection module 606 is configured to perform anomaly detection on the topological nodes using the anomaly detection algorithm, and determine a target topological node corresponding to the abnormal root cause of the target service in the topological nodes according to the detection result.
[0106] In an optional embodiment, the construction module 602 is further configured to: Determining at least two target applications included in the target service, and application relationship data corresponding to the at least two target applications respectively; Determine a graph node based on the at least two target applications, and determine a graph relationship based on application relationship data corresponding to the at least two target applications respectively; The topological graph is constructed by using the graph nodes and the graph relations as the topological data.
[0107] In an optional embodiment, the construction module 602 is further configured to: Determining at least two topological nodes in the topological graph; The indicator data and log data respectively corresponding to the at least two topological nodes are used as the node data.
[0108] In an optional embodiment, the determining module 604 is further configured to: Performing time series signal analysis on the node data; In the case where it is determined according to the analysis result that the node data includes missing data, determining that the data characteristic is data missing; When it is determined according to the analysis result that the node data does not contain missing data, signal periodic processing is performed on the node data, and the data characteristic of the node data is determined according to the processing result.
[0109] In an optional embodiment, the determining module 604 is further configured to: Performing signal period detection on the node data, and when it is determined that the node data corresponds to a periodic signal, performing periodic signal noise reduction on the node data to obtain target node data; In a case where the target node data corresponds to a periodic signal, determining that the data characteristic of the node data is a periodic signal; In a case where the target node data does not correspond to a periodic signal, it is determined that the data characteristic of the node data is a non-periodic signal.
[0110] In an optional embodiment, the determining module 604 is further configured to: In the case where the data characteristic is data missing or a periodic signal, determining a feature detection algorithm corresponding to the topological node, and using the feature detection algorithm as the anomaly detection algorithm; In the case where the data characteristic is a non-periodic signal, a baseline detection algorithm corresponding to the topological node is determined, and the baseline detection algorithm is used as the anomaly detection algorithm.
[0111] In an optional embodiment, the determining module 604 is further configured to: An anomaly type corresponding to the topological node is determined based on the data characteristic, and the anomaly detection algorithm is selected based on the anomaly type.
[0112] In an optional embodiment, the detection module 606 is further configured to: Performing anomaly detection on the topological node using the anomaly detection algorithm to determine anomaly confidence of the topological node; The abnormality confidence is used as the detection result.
[0113] In an optional embodiment, the detection module 606 is further configured to: Determine an abnormal topological node corresponding to a target confidence level among the at least two topological nodes according to the detection result; The abnormal topology node is used as the target topology node corresponding to the abnormal root cause of the target service.
[0114] An anomaly detection device provided by an embodiment of the present specification constructs a topology map based on the topology data of the target service and obtains the node data of the topology node in the topology map. The data characteristics of the node data are analyzed, and the anomaly detection algorithm corresponding to the topology node is determined based on the data characteristics. The anomaly detection algorithm matching the data characteristics is selected according to the data characteristics of the node data, thereby improving the accuracy of anomaly detection for the target service. The anomaly detection algorithm is used to perform anomaly detection on the topology node, and the target topology node corresponding to the abnormal root cause of the target service is determined in the topology node according to the detection result, thereby improving the efficiency of anomaly detection for the target service.
[0115] The above is a schematic scheme of an abnormality detection device of this embodiment. It should be noted that the technical scheme of the abnormality detection device and the technical scheme of the abnormality detection method described above are of the same concept, and the details not described in detail in the technical scheme of the abnormality detection device can be referred to the description of the technical scheme of the abnormality detection method described above.
[0116] Figure 7 The block diagram of a computing device 700 according to an embodiment of the present specification is shown. The components of the computing device 700 include but are not limited to a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and the database 750 is used to store data.
[0117] The computing device 700 also includes an access device 740 that enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0118] In one embodiment of the present specification, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 7 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0119] The computing device 700 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 700 may also be a mobile or stationary server.
[0120] The processor 720 is used to execute the following computer executable instructions, which implement the steps of the above-mentioned abnormality detection method when executed by the processor.
[0121] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned anomaly detection method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above-mentioned anomaly detection method.
[0122] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned abnormality detection method when executed by a processor.
[0123] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned anomaly detection method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned anomaly detection method.
[0124] An embodiment of the present specification also provides a computer program product, including a computer program or instructions, which implement the steps of the above-mentioned abnormality detection method when executed by a processor.
[0125] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the above-mentioned anomaly detection method belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of the above-mentioned anomaly detection method.
[0126] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0127] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0128] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0129] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0130] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. An anomaly detection method, comprising: Construct a topology map based on the topology data of the target service, and obtain node data of the topology nodes in the topology map; Analyzing data characteristics of the node data, and determining an anomaly detection algorithm corresponding to the topological node based on the data characteristics; The anomaly detection algorithm is used to perform anomaly detection on the topological node, and a target topological node corresponding to the abnormal root cause of the target service is determined in the topological node according to the detection result.
2. According to the anomaly detection method of claim 1, the step of constructing a topology map based on the topology data of the target service comprises: Determining at least two target applications included in the target service, and application relationship data corresponding to the at least two target applications respectively; Determine a graph node based on the at least two target applications, and determine a graph relationship based on application relationship data corresponding to the at least two target applications respectively; The topological graph is constructed by using the graph nodes and the graph relations as the topological data.
3. The anomaly detection method according to claim 1, wherein obtaining node data of a topological node in the topological graph comprises: Determining at least two topological nodes in the topological graph; The indicator data and log data respectively corresponding to the at least two topological nodes are used as the node data.
4. The anomaly detection method according to claim 1, wherein analyzing the data characteristics of the node data comprises: Performing time series signal analysis on the node data; In the case where it is determined according to the analysis result that the node data includes missing data, determining that the data characteristic is data missing; When it is determined according to the analysis result that the node data does not contain missing data, signal periodic processing is performed on the node data, and the data characteristic of the node data is determined according to the processing result.
5. The abnormality detection method according to claim 4, wherein the performing signal periodic processing on the node data and determining the data characteristics of the node data according to the processing result comprises: Performing signal period detection on the node data, and when it is determined that the node data corresponds to a periodic signal, performing periodic signal noise reduction on the node data to obtain target node data; In a case where the target node data corresponds to a periodic signal, determining that the data characteristic of the node data is a periodic signal; In a case where the target node data does not correspond to a periodic signal, it is determined that the data characteristic of the node data is a non-periodic signal.
6. The anomaly detection method according to claim 5, wherein the step of determining the anomaly detection algorithm corresponding to the topological node based on the data characteristics comprises: In the case where the data characteristic is data missing or a periodic signal, determining a feature detection algorithm corresponding to the topological node, and using the feature detection algorithm as the anomaly detection algorithm; In the case where the data characteristic is a non-periodic signal, a baseline detection algorithm corresponding to the topological node is determined, and the baseline detection algorithm is used as the anomaly detection algorithm.
7. The anomaly detection method according to claim 1, wherein determining the anomaly detection algorithm corresponding to the topological node based on the data characteristics comprises: An anomaly type corresponding to the topological node is determined based on the data characteristic, and the anomaly detection algorithm is selected based on the anomaly type.
8. The anomaly detection method according to claim 1, wherein the anomaly detection algorithm is used to perform anomaly detection on the topological node, comprising: Performing anomaly detection on the topological node using the anomaly detection algorithm to determine anomaly confidence of the topological node; The abnormality confidence is used as the detection result.
9. The anomaly detection method according to claim 8, wherein when there are at least two topological nodes, determining the target topological node corresponding to the abnormal root cause of the target service in the topological nodes according to the detection result comprises: Determine an abnormal topological node corresponding to a target confidence level among the at least two topological nodes according to the detection result; The abnormal topology node is used as the target topology node corresponding to the abnormal root cause of the target service.
10. An abnormality detection device, comprising: A construction module is configured to construct a topology map based on the topology data of the target service and obtain node data of the topology nodes in the topology map; A determination module, configured to analyze data characteristics of the node data, and determine an anomaly detection algorithm corresponding to the topological node based on the data characteristics; The detection module is configured to perform anomaly detection on the topological nodes using the anomaly detection algorithm, and determine a target topological node corresponding to the abnormal root cause of the target service in the topological nodes according to the detection result.
11. A computing device comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions. When the computer program or instructions are executed by the processor, the steps of the abnormality detection method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium storing a computer program or instruction, wherein the computer program or instruction, when executed by a processor, implements the steps of the anomaly detection method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program or instructions, which, when executed by a processor, implements the steps of the anomaly detection method according to any one of claims 1 to 9.