Abnormality analysis method and device based on CMDB, equipment and storage medium

Through the analysis method of directed graph and exception weighted propagation model based on CMDB, the accuracy, dynamic and automated abnormal analysis problems of business application systems are solved, the analysis efficiency and accuracy are improved, and the resource allocation and response speed are optimized.

CN120455082APending Publication Date: 2025-08-08CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510594160.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art cannot achieve accurate, dynamic and automated abnormal analysis of business application systems, resulting in poor analysis and inefficient efficiency.

Method used

Based on CMDB, we construct a directed graph, and combine the original data and framework description information of multiple target data sources, use the exception weighted propagation model to analyze the abnormal distribution characteristics of the business application system, and realize the dependency analysis of system components and data sources.

Benefits of technology

It realizes the accuracy, dynamic and automated abnormal analysis of business application systems, improves analysis efficiency and accuracy, and optimizes resource allocation and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455082A_ABST
    Figure CN120455082A_ABST
Patent Text Reader

Abstract

The invention discloses a CMDB-based anomaly analysis method, apparatus and device, and a storage medium, and belongs to the field of data processing. The method comprises the steps that fusion data is obtained based on original data collected from multiple target data sources, framework description information of a service application system in a CMDB and multiple data processing layers, and the framework description information is used for representing the dependency relationship between system components in the service application system and the target data sources; according to the framework description information and the fusion data in the CMDB, a directed graph is constructed, the directed graph comprises nodes and edges used for connecting the nodes, the nodes are used for representing a target data source and a system component, and the length of the edges is used for representing the degree of dependence between the connected nodes; and obtaining abnormal distribution characteristic parameters of the business application system through an abnormal weighted propagation model according to the directed graph, the original data and the decision coefficient. According to the embodiment of the invention, the precision requirement of the business application system on anomaly analysis can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a CMDB-based anomaly analysis method, apparatus, device, and storage medium. Background Art

[0002] With the continuous development of internet information technology, the scale and complexity of various business application systems that touch users' lives and work are also increasing. Due to the diverse business types and rapid structural iteration of business application systems, the diversity of anomaly propagation paths is increasing. Anomaly analysis of business application systems based on single indicators, static configuration information, or business anomaly data only provides superficial analysis and fails to provide in-depth analysis. This results in poor anomaly analysis of business application systems and makes it difficult to meet the precision requirements of business application systems for anomaly analysis. Summary of the Invention

[0003] The embodiments of the present application provide a CMDB-based anomaly analysis method, apparatus, device, and storage medium that can meet the business application system's requirements for precise anomaly analysis.

[0004] In a first aspect, an embodiment of the present application provides an anomaly analysis method based on CMDB, comprising: obtaining fused data based on original data collected from multiple target data sources, framework description information of a business application system in a configuration management database (CMDB), and multiple preset data processing layers, wherein the fused data includes integrated data, feature splicing vectors, and decision coefficients, and the framework description information is used to characterize the dependency relationships between system components and target data sources in the business application system; constructing a directed graph based on the framework description information and fused data in the CMDB, wherein the directed graph includes nodes and edges for connecting nodes, wherein the nodes are used to characterize target data sources and system components, and the length of the edges is used to characterize the degree of dependency between connected nodes; obtaining anomaly distribution characteristic parameters of the business application system through an anomaly weighted propagation model based on the directed graph, original data, and decision coefficients.

[0005] In the second aspect, an embodiment of the present application provides an anomaly analysis device based on CMDB, including: a data fusion module, which is used to obtain fused data based on the original data collected from multiple target data sources, the framework description information of the business application system in the configuration management database CMDB and the preset multiple data processing layers, the fused data including integrated data, feature splicing vectors and decision coefficients, and the framework description information is used to characterize the dependency relationship between system components and target data sources in the business application system; a directed graph construction module, which is used to construct a directed graph based on the framework description information and fused data in the CMDB, the directed graph including nodes and edges for connecting nodes, the nodes are used to characterize the target data sources and system components, and the length of the edges is used to characterize the degree of dependency between the connected nodes; an anomaly analysis module, which is used to obtain the anomaly distribution characteristic parameters of the business application system through an anomaly weighted propagation model based on the directed graph, original data and decision coefficients.

[0006] In a third aspect, an embodiment of the present application provides a CMDB-based anomaly analysis device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the CMDB-based anomaly analysis method of the first aspect is implemented.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the CMDB-based anomaly analysis method of the first aspect is implemented.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the CMDB-based anomaly analysis method of the first aspect when the computer program is executed by a processor.

[0009] The embodiment of the present application provides a CMDB-based anomaly analysis method, apparatus, device and storage medium, which can fuse the original data collected by multiple target data sources based on the dependency relationship between the system components and target data sources provided by the CMDB to obtain fused data of multi-source heterogeneous data. According to the fused data and the dependency relationship between the system components and target data sources provided by the CMDB, a directed graph that can reflect the dependency relationship and degree of dependency between the system components and target data sources is constructed. The anomaly distribution characteristic data obtained based on the directed graph, the original data and the decision coefficient is obtained by strengthening the interactive analysis between the system components and the target data sources and capturing the propagation and diffusion laws of anomalies in the business application system, thereby enhancing the ability to analyze anomalies, achieving more accurate anomaly analysis and meeting the precision requirements of the business application system for anomaly analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A flowchart of a CMDB-based anomaly analysis method provided in one embodiment of the present application;

[0012] Figure 2 A schematic diagram of an example architecture of a CMDB-based anomaly analysis system provided in an embodiment of the present application;

[0013] Figure 3 A schematic diagram of an example of an exception analysis process based on CMDB provided in an embodiment of the present application;

[0014] Figure 4 A schematic diagram of the structure of a CMDB-based anomaly analysis device provided in one embodiment of the present application;

[0015] Figure 5 A schematic diagram of the structure of a CMDB-based anomaly analysis device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0016] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating examples of the present application. It should be noted that the acquisition, storage, use, processing, etc. of information and data in the embodiments of the present application are authorized by the user or relevant agencies and comply with the relevant provisions of national laws and regulations.

[0017] With the continuous development of Internet information technology, the scale and complexity of various business application systems that affect users' lives and work are also increasing. Due to the diversified business types and rapid structural iteration characteristics of business application systems, the diversity of anomaly propagation paths is increasing. Anomaly analysis of business application systems based on single indicators, static configuration information, or business anomaly data results in isolated analysis of different aspects of data, lacks cross-data source correlation, and ignores the dynamic changes that may occur during the operation of business application systems. Analysis of anomalies is superficial and lacks in-depth analysis, resulting in poor anomaly analysis results for business application systems and a lag in anomaly analysis results. Deploying professional personnel to manually participate in anomaly analysis will also lead to low anomaly analysis efficiency, exacerbating the lag in anomaly analysis results, making it difficult to meet the business application system's requirements for precision, dynamism, and automation.

[0018] The present application provides an anomaly analysis method, apparatus, equipment, storage medium, and program product based on a Configuration Management Database (CMDB), which can collect data from multiple data sources and, in combination with the framework description information of the business application system in the CMDB, realize the fusion and dynamic analysis of multi-source heterogeneous data, explore the dependency relationships between nodes at all levels in the business application system, as well as the dependency relationships between the data sources of the business application system and nodes at all levels, deepen the analysis of the interaction structure between system components and data sources, capture the diffusion pattern of anomalies, improve the breadth and depth of anomaly analysis, and meet the business application system's requirements for accuracy, dynamism, and automation of anomaly analysis.

[0019] The following describes the CMDB-based anomaly analysis method, device, equipment, storage medium, and program product provided in this application.

[0020] The present application provides a CMDB-based anomaly analysis method, which can be applied to scenarios where anomaly analysis is performed on business application systems. The CMDB-based anomaly analysis method can be executed by CMDB-based anomaly analysis devices, equipment, etc., and is not limited here. Figure 1 A flowchart of an abnormality analysis method based on CMDB provided in one embodiment of the present application is shown in FIG. Figure 1 As shown, the CMDB-based anomaly analysis method may include steps S101 to S103.

[0021] In step S101 , fused data is obtained based on original data collected from multiple target data sources, framework description information of the business application system in the CMDB, and multiple preset data processing layers.

[0022] Business application systems are the analysis targets for anomaly analysis in embodiments of this application. Embodiments of this application can perform anomaly analysis on one or more business application systems. When performing anomaly analysis on two or more business application systems, interactive analysis of system components across these systems can be implemented. A business application system may include multiple system components, which may include but are not limited to functional modules, servers, and databases. Servers may include but are not limited to physical servers and virtual servers. Databases may include information such as database units and database types.

[0023] The target data source includes the selected data source for collecting raw data. For example, the target data source may include but is not limited to a performance monitoring data source, a security monitoring data source, a log collection data source, etc.; performance monitoring data may be obtained from the performance monitoring data source, and the performance monitoring data may include but is not limited to server load, memory usage, etc.; security monitoring data may be obtained from the security monitoring data source, and the security monitoring data may include but is not limited to the number of intrusion prevention system (IPS) attacks, the number of abnormal port communications, the number of data exports during unusual time periods, etc.; log collection data may be obtained from the log collection data source, and the log collection data may include but is not limited to transaction volume, fluctuations in transaction success rate, etc. The raw data collected from the target data source may include historical data and / or real-time data. In some examples, simulation data for different scenarios can be generated based on historical data, and the simulation data can be compared with real-time data. The process of constructing a directed graph and the process of calculating abnormal distribution characteristic parameters are continuously optimized, thereby building an abnormal analysis framework covering the entire business chain of the business application system.

[0024] A CMDB is a logical database that can contain information on the entire lifecycle of configuration items and the relationships between them. These relationships include, but are not limited to, physical and communication relationships. The framework description information in the CMDB can represent the dependencies between system components and target data sources in a business application system. These dependencies include, but are not limited to, dependencies between system components, dependencies between system components and target data sources, and dependencies between target data sources. Different data processing layers have different functions and process the raw data collected from target data sources differently. The data output from different data processing layers also differs. The data output from multiple data processing layers can form fused data, which can be viewed as a unified dataset formed by layered integration of multi-source data. Fused data includes integrated data, feature concatenation vectors, and decision coefficients. In some examples, the multiple data processing layers may include a data integration layer, a feature concatenation layer, and a decision processing layer. The output of the data integration layer includes integrated data, the output of the feature concatenation layer includes feature concatenation vectors, and the output of the decision processing layer includes decision coefficients. The integrated data can be a collection of raw data corresponding to system components extracted from each target data source. In some examples, each system component can correspond to one or more integrated data sets. The integrated data of a system component includes the raw data corresponding to the system component from each target data source. The specific integration method is not limited here. The feature concatenation vector can be implemented as a vector concatenated from feature vectors extracted from the raw data. In some examples, the concatenation order and concatenation method of the feature vectors extracted from the raw data can be determined based on the dependency relationships of the system components represented by the framework description information. The decision coefficient can represent the influence of the system components and the target data source on the propagation of anomalies. In some examples, the decision coefficient can be calculated based on the dependency relationships of the system components guaranteed by the framework description information, the raw data corresponding to the system components in the target data source, etc., or it can be set based on specific scenarios, requirements, experience, etc. A systematic processing method with multiple data processing layers enables the collaborative calculation of data from different fields, which can improve the accuracy and depth of anomaly analysis. Fusion not only breaks down data silos but also reveals potential interactions between different data sources. For example, performance anomalies (i.e., anomalies in a performance data source) can lead to business interruptions (i.e., anomalies in a business data source), while security incidents (i.e., anomalies in a security data source) can lead to performance issues (i.e., anomalies in a security data source). The fusion of multi-source heterogeneous data significantly enhances the breadth and depth of anomaly analysis, revealing cross-dimensional anomaly correlations and enabling more accurate anomaly analysis.

[0025] In step S102, a directed graph is constructed based on the framework description information and fusion data in the CMDB.

[0026] A directed graph includes nodes and edges connecting the nodes. Nodes are used to represent target data sources and system components. System components and target data sources can be used as nodes. The edge connecting two nodes is used to represent the dependency relationship between the two nodes. It should be noted that the edges in a directed graph are pointed. Based on the direction of the edge between the two nodes, the upstream node and the downstream node of the two nodes can be determined. The length of the edge is used to represent the degree of dependency between the connected nodes. In some examples, the shorter the edge length, the higher the degree of dependency between the nodes, and the longer the edge length, the lower the degree of dependency between the nodes. A directed graph can reflect the dependency relationship and degree of dependency between nodes.

[0027] Based on the dependencies between system components and target data sources in a business application system, as represented by the framework description information in the CMDB, it is possible to determine which nodes have edges and the direction of these edges. The length of the edges between these nodes is calculated based on the fused data associated with these nodes. Edges are used to connect these nodes to form a directed graph. This directed graph can demonstrate the degree of dependency between system components in a business application system and the services of the corresponding data source, as well as the impact path of anomalies within the business application system.

[0028] In step S103, the abnormal distribution characteristic parameters of the business application system are obtained through an abnormal weighted propagation model according to the directed graph, the original data and the decision coefficient.

[0029] Based on the raw data and decision coefficient corresponding to a node, the node's anomaly parameter can be obtained. The node's anomaly parameter can reflect the likelihood of an anomaly occurring at the node. In some examples, the node's anomaly parameter is positively correlated with the likelihood of an anomaly occurring at the node; that is, the larger the node's anomaly parameter, the higher the likelihood of an anomaly occurring at the node. The directed graph can be used to determine the upstream and downstream relationships between nodes. The weighted anomaly propagation model can simulate the spread of anomalies within a business application system. The directed graph and the node's anomaly parameters can be input into the weighted propagation model and repeated through multiple iterations to obtain accurate anomaly parameters for each node. The anomaly distribution characteristic parameter of a business application system can reflect the anomaly distribution characteristics of the entire business application system, specifically the concentration of the anomaly distribution. In some examples, the anomaly distribution characteristic parameter can be positively correlated with the concentration of the anomaly distribution; that is, the larger the anomaly distribution characteristic parameter, the more concentrated the anomaly distribution, while the smaller the anomaly distribution characteristic parameter, the more widespread the anomaly distribution. By combining the anomaly parameters of a single node with the anomaly parameters of multiple nodes in the business application system, the concentration of anomalies in the business application system, i.e., the anomaly distribution characteristic parameter, can be determined. Abnormal characteristic distribution parameters can be used to provide technical support for abnormal management of business application systems, thereby executing abnormal management strategies based on abnormal characteristic distribution parameters, effectively improving the real-time nature of abnormal management and the accuracy of decision-making, and optimizing the overall abnormal management efficiency of business application systems.

[0030] In an embodiment of the present application, the raw data collected by multiple target data sources can be fused based on the dependency relationship between the system components and the target data sources provided by the CMDB to obtain fused data of multi-source heterogeneous data. Based on the fused data and the dependency relationship between the system components and the target data sources provided by the CMDB, a directed graph that can reflect the dependency relationship and degree of dependency between the system components and the target data sources is constructed. The abnormal distribution characteristic data obtained based on the directed graph, the raw data and the decision coefficient is obtained by strengthening the interaction analysis between the system components and the target data sources and capturing the propagation and diffusion laws of the abnormalities in the business application system. This enhances the ability to analyze abnormalities and enables more accurate abnormality analysis to meet the business application system's requirements for accurate abnormality analysis. In addition, when the framework description information in the CMDB is updated and new raw data appears in the target data source, the abnormal distribution characteristic parameters will also change accordingly, so that abnormality analysis can be implemented dynamically, meeting the business application system's requirements for dynamic abnormality analysis. In addition, the abnormality analysis process does not require human participation, and the efficiency of abnormality analysis is greatly improved, meeting the business application system's requirements for automated abnormality analysis. Based on the abnormal distribution characteristic parameters, corresponding measures can be taken to reduce the possibility of abnormalities, and the business application system can be optimized. The resource allocation response speed of the optimized business application system is significantly improved. Through experiments, the allocation time of 39 business application systems has been reduced from the 5 hours required by the CMDB-based abnormality analysis method in the embodiment of the present application to the 1 hour required by the CMDB-based abnormality analysis method in the embodiment of the present application, and the process automation rate can be increased to 80%, comprehensively enhancing the abnormal response efficiency and resource management capabilities of the business application system. By analyzing the interactive structure of the business application system, key nodes and paths prone to abnormalities can be accurately identified. Through experiments, 180 abnormal events were tracked through monitoring logs and fault records over a period of 6 months. The average root cause analysis time of abnormal events can be reduced from the 1 hour required by the CMDB-based abnormality analysis method in the embodiment of the present application to the 10 minutes required by the CMDB-based abnormality analysis method in the embodiment of the present application. The root cause analysis speed can be increased by 83%, significantly reducing the missed detection rate, comprehensively covering the nodes prone to abnormalities in the business system, improving the accuracy of abnormality analysis, optimizing the response mechanism, and improving the robustness and business continuity of the business application system.

[0031] In some embodiments, the multiple data processing layers include a data integration layer, a feature splicing layer, and a decision processing layer. The above step S101 can be specifically refined as follows: the data integration layer integrates the original data associated with the same primary key field according to the primary key field in the framework description information to obtain integrated data, where the primary key field is used to characterize the system components; the feature splicing layer splices the feature vectors extracted from the original data according to the dependency relationship of the system components represented by the framework description information to obtain a feature splicing vector; the decision processing layer generates a decision coefficient according to the degree of influence of the target data source and the framework description information on the anomaly analysis; and the integrated data, the feature splicing vector, and the decision coefficient are used to obtain fused data.

[0032] A primary key field represents a system component and is unique; that is, different system components have different primary key fields. Specifically, the primary key field in the CMDB can be used. The data integration layer uses the primary key field to integrate data corresponding to the same primary key field in each target data source, generating integrated data corresponding to each primary key field, and therefore each system component.

[0033] The feature splicing layer can generate association features between system components based on the dependency relationships of system components represented by the framework description information in the CMDB. Then, based on the association features between system components and the original data, it extracts the feature vectors of the original data from multiple target data sources for splicing. It can also perform dimensionality reduction processing on the vectors obtained after splicing to compress redundant information and obtain feature splicing vectors.

[0034] The decision processing layer can determine the degree of influence of the characteristics of the target data source itself and the system components represented by the framework description information and the dependencies of the target data source on the anomaly analysis, and then assign decision coefficients to the system components and the target data source based on the degree of influence. The decision coefficient is used to represent the importance of the target data source and the system components and the uncertainty of the anomaly analysis prediction. The decision processing layer can generate the decision coefficient based on the original data and framework description information of the target data source, or generate the decision coefficient based on the integrated data and the feature splicing vector. In some examples, the decision system may include the weight of the target data source and the weight of the system component, and may also include the influence probability of the target data source and the influence probability of the system component; the weight of the target data source can represent the importance of the target data source, and the weight of the system component can represent the importance of the system component; the influence probability of the target data source can represent the probability of the target data source's influence on the anomaly analysis prediction, and the influence probability of the system component can represent the probability of the system component's influence on the anomaly analysis prediction.

[0035] The collection of integrated data, feature splicing vectors, and decision coefficients can be used as fused data, or they can be processed to form fused data, without limitation. By integrating the three data processing layers—the data integration layer, the feature splicing layer, and the decision processing layer—a closed loop can be achieved, from data to features to decisions.

[0036] In some embodiments, the above step S101 can be specifically refined as follows: generating a dependency matrix based on the framework description information; generating nodes and edges based on the dependency matrix; obtaining the length of the edge based on the fused data; and constructing a directed graph based on the nodes, edges, and edge lengths.

[0037] The dependency relationships between system components and target data sources in the business application system represented by the framework description information can be processed to obtain a dependency matrix. The dependency matrix represents the dependency relationships between the functional modules and functional resources of the business application system. The functional modules include some system components, and the functional resources include some system components and target data sources. That is, some of the system components serve as functional modules, and some serve as functional resources; the target data source serves as a functional resource. The dependency matrix can represent the upstream and downstream relationships between the functional resources corresponding to each functional module. Through the dependency matrix, the degree of dependence of each functional module in the business application system on the functional resources can be obtained, from which the impact path of the failure of the functional resources on the overall business application system can be determined. For example, the business application system is a transaction payment application system. The dependency matrix obtained based on CMDB can be shown in Table 1 below:

[0038] Table 1

[0039]

[0040]

[0041] Among them, the functional modules include the core transaction module, payment gateway module, report generation module, user management module, security audit module, log processing module and traffic analysis module, and the functional resources include physical servers, virtual servers, database types, database units, performance monitoring data sources, security monitoring data sources, and log collection data sources. The physical servers, virtual servers, database types, and database units in the functional resources are system components, and the performance monitoring data sources, security monitoring data sources, and log collection data sources in the functional resources are target data sources. The target data source with a "√" corresponding to the functional module indicates that the functional module covers the target data source with a "√", and the functional module has a dependency relationship with the target data source with a "√". Based on the row of functional resources corresponding to the functional module, the dependency relationship, that is, the upstream and downstream relationship, between the functional module and the corresponding functional resource can be obtained. For example, the core transaction module is implemented by the physical server Server01, the virtual server VM01, and the database unit Unit01 with the database type of MySQL, and the core transaction module can cover the performance monitoring data source, security monitoring data source, and log collection data source. Among them, the core transaction module can be the upstream component of the physical server Server01, the physical server Server01 is the upstream component of the virtual server VM01, the virtual server VM01 is the upstream component of the database MySQL, the database MySQL is the upstream component of the database unit Unit01, and the core transaction module is the upstream component of the performance monitoring data source, the security monitoring data source, and the log collection data source. The dependencies between other performance modules and performance resources are not detailed here. The dependency matrix can be used to determine the impact path of the failure of functional resources on the entire business application system. For example, if the database MySQL fails, it will affect the virtual server VM01, virtual server VM05, virtual server VM06, physical server Server01, physical server Server04, physical server Server02, database unit Unit01, database unit Unit05, database unit Unit06, core transaction module, security audit module and log processing module.

[0042] Functional modules and functional resources in the dependency matrix can be identified as nodes. Specifically, system components and target data sources in the dependency matrix can be identified as nodes. Based on the dependencies between the system components and target data sources reflected in the dependency matrix, the positions and directions of the edges connecting the nodes are determined. Two nodes connected by an edge have a dependency relationship, with the edge pointing from the upstream node to the downstream node. Based on the direction of the edge between the nodes, the upstream and downstream nodes of each node can be determined. The edge length can also be calculated based on the fused data corresponding to the nodes connected by the edge. For example, the fused data includes integrated data, feature concatenation vectors, and a decision coefficient. The feature concatenation vectors and integrated data can be weighted and fused according to the decision coefficient to determine the edge length. Alternatively, the edge length can be determined based on the similarity between the node feature concatenation vectors and integrated data, combined with the decision coefficient. The edge length can reflect the degree of dependency between two nodes in a dependency relationship. In some examples, the shorter the edge length, the higher the degree of dependency between the two nodes connected by the edge. After determining the nodes, edges, and edge lengths, the corresponding nodes can be connected using the edges to construct the directed graph.

[0043] In some embodiments, key nodes for anomaly analysis can be screened in a directed graph, so that subsequent anomaly analysis can be performed with the key nodes as the core to obtain anomaly distribution characteristic parameters for the business application system. Specifically, the shortest distance between two nodes connected by an edge can be calculated based on the length of an edge in the directed graph. For any node, the centrality parameter of any node can be obtained by using the shortest distance between any node and other nodes connected by edges in the directed graph and the number of nodes in the directed graph. Nodes whose centrality parameters meet preset conditions are selected as key nodes for anomaly analysis.

[0044] The distance between two nodes can be determined based on the length of the edge directly connecting the two nodes, or based on the sum of the lengths of the edges connecting the two nodes to an intermediate node. It should be noted that the length of the edge directly connecting the two nodes may be different from the sum of the lengths of the edges connecting the two nodes to an intermediate node. For example, node A1 is connected to node A2 via edge B1, node A1 is connected to node A3 via edge B2, and node A2 is connected to node A3 via edge B3. The distance between node A1 and node A3 may be the length of edge B2, or the sum of the length of edge B1 and the length of edge B3. However, the length of edge B2 and the sum of the length of edge B1 and the length of edge B3 may be different. The shortest distance between two nodes can be selected from the sum of the lengths of the edges directly connecting the two nodes and the sum of the lengths of the edges connecting the two nodes to an intermediate node. For example, the shortest distance between node v and node u can be obtained by the following formula (1):

[0045] d′(v,u)=min(d(v,u),d(v,w)+d(w,u)) (1)

[0046] Among them, d′(v,u) is the shortest distance between node v and node u; d(v,u) is the length of the edge connecting node v and node u; d(v,w) is the length of the edge connecting node v and node w; d(w,u) is the length of the edge connecting node w and node u; min() is the minimum value algorithm.

[0047] The shortest distance between two nodes can be used to obtain the dependency path with the highest degree of dependency between the two nodes. This dependency path can also be considered as the propagation path along which the anomaly will propagate when an anomaly occurs. Key nodes can exhibit characteristics such as high connectivity, high anomaly propagation, and susceptibility to single-point failures. The importance of a node in a directed graph can be measured by its close centrality. The total distance from a node to other nodes in the directed graph can reflect the close centrality of the node. The smaller the total distance from a node to other nodes in the directed graph, the closer the node is to the core position of the node network in the directed graph. Close centrality can be quantified by a centrality parameter. For example, the centrality parameter of a node can be obtained according to the following formula (2):

[0048]

[0049] Among them, G′ closeness (v) is the centrality parameter of node v; N is the total number of nodes in the directed graph; ∑ u∈V d(v,u) is the sum of the distances from node v to all other nodes in the directed graph.

[0050] The preset conditions for determining key nodes can be set based on scenarios, needs, experience, and the like, and are not limited here. For example, the preset conditions may include selecting the top X nodes ranked by centrality parameter from highest to lowest as key nodes, or the preset conditions may include selecting nodes with centrality parameters greater than or equal to a preset parameter threshold as key nodes, but the preset conditions are not limited thereto. After obtaining the key nodes, in the process of calculating the abnormal distribution characteristic parameters of the business application system, multiple iterative operations can be performed based on the key nodes, raw data, and decision coefficients in the directed graph to obtain the abnormal distribution characteristic parameters of the business application system. That is, the abnormal distribution characteristic parameters are calculated based on the dependencies of the key nodes, the raw data associated with the key nodes, and the decision coefficients. This calculation of the abnormal distribution characteristic parameters of the business application system does not require the participation of all nodes in the directed graph, thereby saving computing resources and further improving the efficiency of anomaly analysis. It can also, to a certain extent, avoid the influence of nodes that have no effect on the anomaly analysis, reduce deviations in the anomaly analysis process, and further improve the accuracy of anomaly analysis. Identifying key nodes and focusing on monitoring them can also effectively reduce the possibility of business terminal and data leakage.

[0051] It should be noted that the above process of determining key nodes and calculating anomaly distribution characteristic parameters is dynamic, tracking the dynamic changes in the dependencies between nodes in the directed graph and updating the key nodes and anomaly distribution characteristic parameters. The results of each shortest path calculation and key node identification can be stored to intuitively understand the dynamic changes of key nodes and better observe the dynamic changes of anomaly analysis.

[0052] In some embodiments, the above-mentioned step S103 can be specifically refined as follows: for any node in the directed graph, according to the abnormal parameters of any node, the propagation weight of the upstream node of any node, the abnormal parameters of the upstream node and the preset attenuation factor, iterative calculation is performed through the weighted propagation model to update the abnormal parameters of any node until the updated abnormal parameters of any node meet the preset convergence conditions; based on the updated abnormal parameters of at least some nodes that meet the preset convergence conditions, the abnormal distribution characteristic parameters are calculated.

[0053] The iterative calculation process of the anomaly parameters of each node and the calculation process of the anomaly distribution characteristic parameters simulate and evaluate how the anomaly propagates among the nodes in the directed graph, and considers the dependency and propagation strength between the nodes, thereby simulating the process of the anomaly spreading from the node to the entire directed graph network. The anomaly parameter of the node can characterize the possibility of an anomaly occurring at the node, and the anomaly parameter can be obtained based on the original data and the decision coefficient. The propagation weight of the upstream node of a node can be obtained from the decision coefficient in the fused data, which can characterize the dependency strength between the upstream node and this node. The attenuation factor can characterize the ratio of the anomaly spreading between nodes to the anomaly retained by the node. The attenuation factor can be set according to the scenario, requirements, experience, etc., and is not limited here. The anomaly parameter of the node can be gradually approached to a stable state through multiple iterative calculations. For example, each iterative calculation can be obtained according to the following formula (3):

[0054]

[0055] in, is the abnormal parameter of node v at time t+1; is the abnormal parameter of node v at the tth time; α is the attenuation factor; N(v) includes all upstream nodes of node v; ω uv is the transmission intensity from node u to node v.

[0056] The preset convergence condition can quantify the convergence process of the abnormal parameters of the node gradually approaching the stable state. Specifically, the maximum change of the abnormal parameters, that is, the maximum change of the abnormal parameters of the node v calculated in two adjacent iterative calculations, can be used as the core factor of the preset convergence condition. In some examples, the preset convergence condition may include that the absolute value of the difference between the abnormal parameters of the node calculated in two adjacent iterative calculations is less than the deviation allowable threshold. The deviation allowable threshold is the maximum value of the allowable deviation of the iterative operation convergence, which can be set according to the scenario, requirements, and experience, and is not limited here. For example, the preset convergence condition can be obtained according to the following formulas (4) and (5):

[0057]

[0058] ΔR<ε (5)

[0059] Wherein, ΔR is the maximum change of the abnormal parameter of node v calculated between two adjacent iterative calculations; ε is the deviation allowable threshold; max is the maximum value algorithm; the description of other parameters can be found in the relevant content of the above embodiment and will not be repeated here.

[0060] The iterative calculation of the anomaly parameters of a node can characterize the propagation behavior of anomalies in a directed graph. The iterative calculation of the anomaly parameters of the above nodes can be achieved by constructing a mathematical model, namely, an anomaly weighted propagation model, to describe the diffusion process of anomalies from one node to adjacent nodes.

[0061] After the abnormal parameters of the nodes converge and stabilize, the abnormal distribution characteristic parameters of the business application system are calculated based on the abnormal parameters of the nodes after convergence and stabilization. At least some of the nodes can be implemented as all nodes in the directed graph, that is, the abnormal parameters of each node in the directed graph can be iteratively calculated, and the abnormal distribution characteristic parameters of the business application system are calculated using the abnormal parameters of each node; or, at least some of the nodes can also be implemented as key nodes in the directed graph, that is, only the abnormal parameters of the key nodes in the directed graph can be iteratively calculated, and the abnormal distribution characteristic parameters of the business application system are calculated using the abnormal parameters of the key nodes. The abnormal distribution characteristic parameters can be used to characterize the degree of abnormal concentration of nodes in the directed graph, and the abnormal distribution characteristic parameters can be obtained by comparing the abnormal parameters of nodes with a high probability of abnormality with the abnormal parameters of the total nodes in the directed graph.

[0062] In some examples, since a node's abnormal parameters gradually increase as the abnormality propagates, normalization can be used to eliminate the accumulated errors in the abnormal parameters as the abnormality propagates, improving the intuitiveness and reliability of the abnormal parameters. Specifically, normalization can be performed on the updated abnormal parameters of at least some nodes that meet a preset convergence condition to obtain a processed abnormal parameter; the sum of the processed abnormal parameters of at least some nodes is determined as the system's global abnormal parameter; and an abnormal distribution characteristic parameter is obtained based on the system's global abnormal parameter and the processed abnormal parameters of the nodes of interest, where the nodes of interest include at least some nodes whose abnormality, as represented by the processed abnormal parameters, meets the preset abnormality condition.

[0063] The abnormal parameters of the nodes can be normalized using the maximum value of the abnormal parameters of the nodes and the minimum value of the abnormal parameters of the nodes in at least some nodes. For example, the abnormal parameters after processing can be obtained according to the following formula (6):

[0064]

[0065] Among them, R′(v) is the abnormal parameter of node v after processing; R(v) is the abnormal parameter of node v before normalization; R max is the maximum value among the abnormal parameters of the nodes in at least some nodes; R min is the minimum value among the abnormal parameters of the nodes in at least some of the nodes.

[0066] The system global abnormality parameter may be the sum of the processed abnormality parameters of at least some nodes. The preset abnormality condition may be used to screen the nodes of interest. In some examples, the preset abnormality condition may include the first M nodes whose abnormality levels are represented by the processed abnormality parameters in descending order, and the ratio of the sum of the processed abnormality parameters of the first M nodes in descending order to the abnormality level may be used as the abnormality distribution characteristic parameter. For example, the abnormality distribution characteristic parameter may be based on the following equations (7) and (8):

[0067] R total =∑ u∈V R′(v) (7)

[0068]

[0069] Among them, R total is the global abnormal parameter of the system; R′(v) is the abnormal parameter of node v after processing; V top is the set of focus nodes; R con is the characteristic parameter of abnormal distribution.

[0070] A higher anomaly distribution characteristic parameter indicates a more concentrated anomaly distribution; a lower anomaly distribution characteristic parameter indicates a more widespread anomaly distribution. The calculation process of the anomaly distribution characteristic parameter quantifies the characteristics of anomaly distribution from a node to a global perspective. The anomaly distribution characteristic parameter supports the identification of nodes of interest and analysis of the concentration of anomaly distribution. It can be applied to scenarios such as continuity management of business application systems, dependency analysis, and system optimization, but is not limited here.

[0071] In some examples, a preset distribution threshold can be used to distinguish between two abnormal distribution situations: concentrated abnormal distribution and widespread abnormal distribution, so as to implement different processing strategies for different abnormal distribution situations. The preset distribution threshold can be set according to the scenario, requirements, experience, etc., and will not be described in detail here. Specifically, after obtaining the abnormal distribution characteristic parameter, when the abnormal distribution characteristic parameter is higher than or equal to the preset distribution threshold, the system component corresponding to the focus node can be subjected to abnormal control processing. The focus node includes the node whose abnormal degree represented by the updated abnormal parameter meets the preset abnormal condition; when the abnormal distribution characteristic parameter is lower than the preset distribution threshold, the business application system as a whole is subjected to robustness improvement processing. When the abnormal distribution characteristic parameter is higher than or equal to the preset distribution threshold, it means that the abnormal distribution in the business application system is concentrated, and the abnormality is mainly distributed in the focus node. Correspondingly, the focus node is subjected to abnormal control processing. The abnormal control processing may include a design that can reinforce or redundancy the focus node. For example, the abnormal control processing may include updating the business application system to a distributed architecture, and / or adding redundant servers, redundant databases and other structures. The specific method of abnormal control processing is not limited here. If the anomaly distribution characteristic parameter is lower than the preset distribution threshold, it indicates that anomalies are widespread in the business application system and that the entire business application system needs to be robustly improved to improve its overall robustness. For example, fault tolerance mechanisms can be added to the business application system and / or the backup and recovery functions of the business application system can be enhanced. The specific method of robustness improvement is not limited here. Anomaly analysis can be used to calculate and discover abnormal changes, and then the anomaly propagation path can be adjusted through anomaly control and robustness improvement processing, thereby forming a closed-loop management of anomaly identification, evaluation, and optimization.

[0072] In some embodiments, before obtaining the fused data, the data source may be screened to determine the target data source, and / or the collected raw data may be preprocessed so that the preprocessed raw data can be used in subsequent processes to obtain the fused data. Specifically, before obtaining the fused data, at least one of the following may be performed: determining the target data source of the business application system based on user application information, the CMDB's business chain, and compliance screening conditions, and collecting raw data from the target data source; and preprocessing the raw data to obtain preprocessed raw data, where the data preprocessing includes at least one of data cleaning, data standardization, and semantic association processing.

[0073] User application information is used to request anomaly analysis of business application systems. This information includes user permission information and business application system information. User permission information indicates the role and permissions of the user initiating the application, enabling the selection of target data sources for which the user has permission. Business application system information identifies the business application system for which the user is requesting anomaly analysis, thereby filtering target data sources based on the business application system. The CMDB's business chain is derived from the dependencies between system components within the CMDB. This chain can be used to filter target data sources that match the business needs of the business application system. Compliance filtering conditions can be set based on requirements to filter target data sources that comply with relevant regulations and requirements within the business application system. By combining user application information, the CMDB's business chain, and compliance filtering conditions, target data sources that meet these requirements can be selected from multiple data sources. This makes the collected raw data more targeted for anomaly analysis of the business application system, further improving the accuracy and relevance of anomaly analysis for the business application system. In some examples, if the user application information conflicts with the business chain and compliance screening conditions of the CMDB, a prompt message may be fed back to the user. The prompt message may contain conflicting information to prompt the user that he or she does not have permission to use one or more data sources, or to prompt the user to change the specific content of the user application information so that the user application information no longer conflicts with the business chain and compliance screening conditions of the CMDB.

[0074] In order to improve data quality and consistency, the raw data is preprocessed. Data cleaning processing may include but is not limited to processing of missing values, outlier values, and duplicate data removal. The integrity and consistency of the raw data after data cleaning processing are improved. Data standardization processing may include but is not limited to processing of unified data formats, unified naming rules, and numerical range adjustment. The comparability and standardization of raw data after data standardization are improved. Semantic association processing is used to establish logical relationships and contextual associations between raw data. By using the logical relationships and contextual associations between fields in the raw data, the business value and association depth of the raw data are enhanced, which can provide more and deeper data for subsequent anomaly analysis.

[0075] In some embodiments, after obtaining the fused data, data quality monitoring can be performed on the fused data. A combination of offline and real-time computing methods can be used to ensure comprehensiveness and efficiency of data quality monitoring. Specifically, the historical data in the fused data can be processed using an offline computing method to obtain historical data quality assessment parameters, and the real-time data in the fused data can be processed using a real-time computing method to obtain real-time data quality assessment parameters. Based on the historical data quality assessment parameters and / or the real-time data quality assessment parameters, fused data that meets preset quality requirements can be screened.

[0076] The offline computing method can be used to process large amounts of historical data in batches, and specifically, an incremental update mechanism can be used for processing, thereby improving processing efficiency. Real-time computing can quickly respond and make decisions through stream processing technology, thereby quickly discovering abnormal real-time data and ensuring the timeliness of quality monitoring. Historical data quality assessment parameters can characterize the data quality of historical data, and real-time data quality assessment parameters can characterize the data quality of real-time data. Quality requirements can be determined based on scenarios, needs, experience, etc. For example, if the historical data quality assessment parameters and real-time data quality assessment parameters are specifically implemented as data missing rate, the quality requirements may include a data missing rate lower than a preset missing rate threshold. The preset missing rate threshold can be set based on scenarios, needs, experience, etc., such as a preset missing rate threshold of 3%. The types of historical data quality assessment parameters and real-time data quality assessment parameters are not limited here, and other parameters that can reflect data quality are also within the scope of protection of the embodiments of this application. Data quality monitoring can standardize data standards, and the fused data that meets the preset quality requirements can meet the needs of subsequent processes such as directed graph construction and calculation of abnormal distribution characteristic parameters.

[0077] In some embodiments, after constructing a directed graph and obtaining the dependency relationships and dependency levels of system components and target data sources, the dependency relationships and dependency levels can be tracked and detected, and potential dependency issues can be identified and addressed in a timely manner, thereby improving the reliability of the business application system. Specifically, the dependency relationships between nodes represented by the directed graph can be tracked and detected to identify dependency anomalies; and anomaly reduction strategies that match the dependency anomalies are executed. A directed graph can represent the dependency relationships and dependency levels between nodes. Through the directed graph, obvious dependency anomalies can be identified. Dependency anomalies include anomalies that can be directly determined from dependency relationships. For example, dependency anomalies may include, but are not limited to, the existence of unreasonable dependency chains, redundant dependencies, single points of failure, performance bottlenecks on a single node, and lack of redundant dependencies. A correspondence between dependency anomalies and anomaly reduction strategies can be pre-set, and an anomaly reduction strategy that matches the identified dependency anomaly is determined from the correspondence. The anomaly reduction strategy is executed to reduce the likelihood of dependency anomalies. For example, if the dependency anomaly includes an unreasonable dependency chain, the anomaly reduction strategy may include decoupling direct dependencies and changing the unreasonable dependency chain to a reasonable dependency chain; if the dependency anomaly includes redundant dependencies, the anomaly reduction strategy may include removing duplicate or unused dependencies; if the dependency anomaly includes a single point of failure, the anomaly reduction strategy may include deploying multiple instances or adopting a master-slave structure; if the dependency anomaly includes a performance bottleneck in a single node, the anomaly reduction strategy may include horizontally scaling the node or optimizing the single node; if the dependency anomaly includes a lack of redundant dependencies, the anomaly reduction strategy may include configuring a backup node for the node. By tracking and detecting the dependency relationships between nodes, potential deviations can be discovered and corrected in a timely manner, providing a traceable anomaly propagation path, providing support for anomaly analysis, and supporting improved stability and business continuity of business application systems.

[0078] In some embodiments, a CMDB verification mechanism can also be integrated, even if data deviations in the CMDB are discovered. Specifically, the current framework structure of the business application system can be determined based on the actual operating data of the business application system during operation; and the framework description information of the CMDB can be updated based on the comparison results between the current framework structure and the framework description information of the CMDB. The functional interface in the business application system can be obtained by obtaining the actual operating data from the business application system, and the current framework structure of the business application system can be determined based on the data flow of the functional interface. If the current framework structure of the determined business application system is consistent with the framework structure represented by the framework description information of the CMDB, there is no need to update the framework description information of the CMDB; if the current framework structure of the determined business application system is inconsistent with the framework structure represented by the framework description information of the CMDB, the framework description information of the CMDB is updated to make the framework structure represented by the framework description information of the CMDB consistent with the actual framework structure of the business application system. Through the CMDB verification mechanism, not only can the data deviations of the CMDB be discovered in a timely manner, but the CMDB can also be updated in reverse.

[0079] The CMDB-based anomaly analysis method in the embodiment of the present application can be implemented on the CMDB-based anomaly analysis system architecture. Figure 2 The schematic diagram of the architecture of an example of an abnormality analysis system based on CMDB provided in the embodiment of the present application is as follows: Figure 2 As shown, the CMDB-based anomaly analysis system may include a data collection layer 21 , a data fusion layer 22 , a dependency modeling layer 23 and an anomaly propagation algorithm layer 24 .

[0080] The data collection layer 21 may have multiple collectors for collecting data from different data sources. Different collectors may be adopted or set according to the needs and application scenarios of the CMDB-based anomaly analysis system. CMDB may provide system components of the business application system and structured information of the system components, so that data of each dimension can be fully collected according to the business chain. The data collection layer 21 may include a database collector 211, an external service data collector 212, and a log collector 213. The database collector 211 supports various database connections and different data query methods, and is used to collect data from the database in real time or in batches. The external service data collector 212 is used to exchange data with external services, cloud platforms, or third-party systems through the Application Programming Interface (API), and completes data collection through authentication, calling, and data processing. The log collector 213 is used to extract log information from log files or log management systems such as Elasticsearch (ES), and can support multiple log formats such as JSON, XML, and Plain Text. In order to cope with sudden increases in traffic, a log collector 213 with high concurrent processing capabilities can also be used, and a dynamic adjustment strategy can be adopted to automatically adjust the collection rate according to the source system load to ensure efficient and stable log collection.

[0081] The data fusion layer 22 can rely on the primary key fields in the CMDB to uniformly map multi-source heterogeneous data, integrate data from multiple heterogeneous data sources, and merge them into a unified data structure for cross-domain analysis and in-depth mining. The data fusion layer 22 includes a data preprocessing unit 221, a data fusion unit 222, and a quality monitoring unit 223. The data preprocessing unit 221 can be used to clean, standardize, and semantically associate the collected data to ensure the consistency, accuracy, and availability of the data, providing a data foundation for subsequent data fusion. The data fusion unit 222 can achieve a full-link closed loop from data to features to decisions through the fusion of three levels: the data integration layer, the feature splicing layer, and the decision processing layer. The quality monitoring unit 223 combines offline computing with real-time computing to ensure that the high-quality data outputted in the end can support abnormal analysis and decision-making.

[0082] The dependency modeling layer 23 can model multi-level dependencies between system components based on the CMDB. This layer can include a multi-level dependency modeling unit 231 and a dependency tracking and detection unit 232. Based on the CMDB, the multi-level dependency modeling unit 231 can visualize and analyze anomalies within and across layers of dependencies by layering, decomposing, and modeling business application systems. The dependency tracking and detection unit 232 can monitor, record, and verify dependencies between system components, services, data, or resources, establishing a clear dependency tracking and detection mechanism.

[0083] The exception propagation algorithm layer 24 can calculate the path of exception propagation and assess its impact based on CMDB dependency modeling, accurately identifying the exception propagation path and quantifying its impact. The exception propagation algorithm layer 24 may include an exception propagation path calculation unit 241 and an exception impact assessment unit 242. The exception propagation path calculation unit 241 is used to determine the propagation path of an exception and identify key nodes and exception propagation links. The exception impact assessment unit 242 can focus on quantifying the actual impact of exception propagation on business application systems or businesses, providing support for exception management and decision-making.

[0084] Based on the above-mentioned CMDB-based anomaly analysis system, an example is used here to illustrate the CMDB-based anomaly analysis process in the embodiment of the present application. Figure 3 A schematic diagram of an example of an abnormality analysis process based on CMDB provided in an embodiment of the present application is shown as follows: Figure 3 As shown in FIG, the CMDB-based anomaly analysis process may include steps a1 to a13. Steps a1 to a5 are performed by the data acquisition layer, steps a6 to a8 are performed by the data fusion layer, steps a9 to a11 are performed by the dependency modeling layer, and steps a12 to a13 are performed by the anomaly propagation algorithm layer.

[0085] In step a1, the data source is identified. This can be achieved by determining whether the data source is available. If so, step a2 is executed; if not, step a3 is executed. This determination can be made based on the business application system information in the user's application and the business chain in the CMDB.

[0086] In step a2, the data source's permissions and compliance are checked. This can be accomplished by determining whether the permissions and compliance requirements are met. If so, step a4 is executed; if not, step a5 is executed. This data source permissions and compliance check can be performed based on the user's permission information and compliance screening criteria in the user's application.

[0087] In step a3, the data collection plan is adjusted to make the data source available.

[0088] In step a4, raw data is collected.

[0089] In step a5, adjust the data collection plan to make the data source consistent with permissions and compliance.

[0090] In step a6, data preprocessing is performed, which specifically includes data cleaning, data standardization, and semantic association.

[0091] In step a7, data fusion is performed. Data fusion specifically includes processing at the data integration layer, processing at the feature splicing layer, and processing at the decision processing layer.

[0092] In step a8, data quality monitoring is performed. This can be accomplished by determining whether the data quality meets quality requirements; if so, executing step a9; if not, executing step a10. Data quality monitoring can be performed using offline real-time computation, which will not be detailed here.

[0093] In step a9, dependency modeling is performed, which can be implemented by constructing a directed graph.

[0094] In step a10, a notification message is sent, which is used to prompt that the data quality does not meet the quality requirements.

[0095] In step a11, dependency relationships are tracked and detected.

[0096] In step a12, anomaly propagation calculation is performed. The anomaly propagation calculation can be specifically implemented by iteratively calculating and determining the anomaly parameters of the node.

[0097] In step a13, the abnormal impact assessment is performed. The abnormal impact assessment can be specifically implemented by calculating abnormal distribution characteristic parameters.

[0098] The specific contents of the above steps a1 to a13 can be found in the relevant descriptions in the above embodiments, which will not be repeated here.

[0099] This application also provides an abnormality analysis device based on CMDB. Figure 4 A schematic diagram of the structure of an abnormality analysis device based on CMDB provided in one embodiment of the present application is shown as follows: Figure 4 As shown, the CMDB-based anomaly analysis device 300 may include a data fusion module 301 , a directed graph construction module 302 , and an anomaly analysis module 303 .

[0100] The data fusion module 301 can be used to obtain fused data based on the original data collected from multiple target data sources, the framework description information of the business application system in the configuration management database CMDB, and the preset multiple data processing layers. The fused data includes integrated data, feature splicing vectors and decision coefficients. The framework description information is used to characterize the dependency relationship between system components and target data sources in the business application system.

[0101] The directed graph construction module 302 can be used to construct a directed graph based on the framework description information and fusion data in the CMDB. The directed graph includes nodes and edges for connecting nodes. The nodes are used to represent target data sources and system components, and the length of the edges is used to represent the degree of dependency between the connected nodes.

[0102] The anomaly analysis module 303 may be used to obtain anomaly distribution characteristic parameters of the business application system through an anomaly weighted propagation model based on the directed graph, original data, and decision coefficients.

[0103] In some embodiments, the multiple data processing layers include a data integration layer, a feature stitching layer, and a decision processing layer.

[0104] The data fusion module 301 can be specifically used to: integrate the original data associated with the same primary key field according to the primary key field in the framework description information through the data integration layer to obtain integrated data, and the primary key field is used to characterize the system components; splice the feature vectors extracted from the original data according to the dependency relationship of the system components represented by the framework description information through the feature splicing layer to obtain a feature splicing vector; generate a decision coefficient according to the influence of the target data source and the framework description information on the anomaly analysis through the decision processing layer, and the decision coefficient is used to characterize the importance of the target data source and the system components and the uncertainty of the anomaly analysis prediction; obtain fused data based on the integrated data, the feature splicing vector and the decision coefficient.

[0105] In some embodiments, the directed graph construction module 302 can be specifically used to: generate a dependency matrix based on the framework description information, the dependency matrix represents the dependency relationship between the functional modules and functional resources of the business application system, the functional modules include some system components, and the functional resources include some system components and target data sources; generate nodes and edges based on the dependency matrix, and there is a dependency relationship between two nodes connected by an edge; obtain the length of the edge based on the fused data; and construct a directed graph based on the nodes, edges and the length of the edges.

[0106] In some embodiments, the CMDB-based anomaly analysis device 300 may further include a key node determination module. The key node determination module may be configured to: calculate the shortest distance between two nodes connected by an edge in a directed graph based on the length of an edge; obtain a centrality parameter for any node using the shortest distance between any node and any other node connected by an edge in the directed graph and the number of nodes in the directed graph; and select nodes whose centrality parameters meet preset conditions as key nodes for anomaly analysis.

[0107] The anomaly analysis module 303 can be specifically used to perform multiple iterative operations based on key nodes, original data and decision coefficients in the directed graph to obtain anomaly distribution characteristic parameters of the business application system.

[0108] In some embodiments, the anomaly analysis module 303 can be specifically used to: for any node in the directed graph, based on the anomaly parameters of any node, the propagation weight of the upstream node of any node, the anomaly parameters of the upstream node and the preset attenuation factor, perform iterative calculations through a weighted propagation model to update the anomaly parameters of any node until the updated anomaly parameters of any node meet the preset convergence conditions, and the anomaly parameters are obtained based on the original data and the decision coefficient; based on the updated anomaly parameters of at least some nodes that meet the preset convergence conditions, calculate the anomaly distribution characteristic parameters, and the anomaly distribution characteristic parameters are used to characterize the degree of anomaly concentration of nodes in the directed graph.

[0109] In some examples, the anomaly analysis module 303 can be specifically used to: normalize the updated anomaly parameters of at least some nodes that meet preset convergence conditions to obtain processed anomaly parameters; determine the sum of the processed anomaly parameters of at least some nodes as the system global anomaly parameters; obtain anomaly distribution characteristic parameters based on the system global anomaly parameters and the processed anomaly parameters of the focus nodes, where the focus nodes include nodes in at least some nodes whose degree of anomaly represented by the processed anomaly parameters meets the preset anomaly conditions.

[0110] In some embodiments, the CMDB-based anomaly analysis device 300 may further include an execution module. The execution module may be configured to: when the anomaly distribution characteristic parameter is greater than or equal to a preset distribution threshold, perform anomaly control processing on the system components corresponding to the focus node, where the focus node includes a node whose anomaly degree represented by the updated anomaly parameter meets the preset anomaly condition; and when the anomaly distribution characteristic parameter is lower than the preset distribution threshold, perform robustness improvement processing on the entire business application system.

[0111] In some embodiments, the CMDB-based anomaly analysis device 300 may further include a data acquisition module and / or a pre-processing module.

[0112] The data collection module can be used to determine the target data source of the business application system based on user application information, the CMDB's business chain, and compliance screening conditions, and collect raw data from the target data source. User application information includes user permission information and business application system information.

[0113] The preprocessing module can be used to perform data preprocessing on the original data to obtain preprocessed original data. The data preprocessing includes at least one of data cleaning processing, data standardization processing, and semantic association processing. The semantic association processing is used to establish logical relationships and contextual associations between the original data.

[0114] In some embodiments, the CMDB-based anomaly analysis device 300 may further include a quality assessment module. The quality assessment module may be configured to: process historical data in the fused data using an offline computing method to obtain historical data quality assessment parameters; process real-time data in the fused data using a real-time computing method to obtain real-time data quality assessment parameters; and screen the fused data to obtain data that meets preset quality requirements based on the historical data quality assessment parameters and / or the real-time data quality assessment parameters.

[0115] In some embodiments, the CMDB-based anomaly analysis device 300 may further include a tracking module that can be used to: track and detect dependency relationships between nodes represented by a directed graph, identify dependency anomaly events, and execute anomaly reduction strategies that match the dependency anomaly events.

[0116] In some embodiments, the CMDB-based anomaly analysis device 300 may further include a reverse update module. The reverse update module may be configured to: determine the current framework structure of the business application system based on actual operational data of the business application system during operation; and update the framework description information in the CMDB based on a comparison result between the current framework structure and the framework description information in the CMDB.

[0117] It should be noted that the CMDB-based anomaly analysis device 300 is a device corresponding to the above-mentioned CMDB-based anomaly analysis method. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of the device and can achieve the same technical effects.

[0118] This application also provides an abnormality analysis device based on CMDB. Figure 5 A schematic diagram of the structure of an abnormality analysis device based on CMDB provided in one embodiment of the present application is shown as follows: Figure 5 As shown, the CMDB-based anomaly analysis device 400 includes a memory 401 , a processor 402 , and a computer program stored in the memory 401 and executable on the processor 402 .

[0119] In some examples, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0120] The memory 401 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the CMDB-based anomaly analysis method according to the embodiments of the present application.

[0121] The processor 402 reads the executable program code stored in the memory 401 to run a computer program corresponding to the executable program code, so as to implement the CMDB-based anomaly analysis method in the above embodiment.

[0122] In some examples, the CMDB-based anomaly analysis device 400 may further include a communication interface 403 and a bus 404. Figure 5 As shown, the memory 401 , the processor 402 , and the communication interface 403 are connected via a bus 404 and communicate with each other.

[0123] The communication interface 403 is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. Input devices and / or output devices can also be connected through the communication interface 403.

[0124] The bus 404 includes hardware, software, or both, coupling the components of the CMDB-based anomaly analysis device 400 to each other. By way of example and not limitation, the bus 404 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 404 may include one or more buses. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.

[0125] The present application also provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the CMDB-based anomaly analysis method in the above-mentioned embodiment can be implemented, and the same technical effect can be achieved. To avoid repetition, the above-mentioned computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which is not limited here.

[0126] The present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the CMDB-based anomaly analysis method in the above embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0127] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. For device embodiments, equipment embodiments, and computer-readable storage medium embodiments, the relevant parts can be referred to the description part of the method embodiment. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of this application. In addition, for the sake of brevity, a detailed description of known method technologies is omitted here.

[0128] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or the flowchart and the combination of the boxes in the block diagram and / or the flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.

[0129] Those skilled in the art should understand that the above embodiments are illustrative rather than restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, the specification and the claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other devices or steps; the quantifier "one" does not exclude a plurality; the terms "first" and "second" are used to identify names rather than to indicate any specific order. Any figure marks in the claims should not be understood as limiting the scope of protection. The functions of multiple parts appearing in the claims can be implemented by a separate hardware or software module. The fact that certain technical features appear in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.

Claims

1. A CMDB-based anomaly analysis method, characterized in that: include: Based on the raw data collected from multiple target data sources, the framework description information of the business application system in the configuration management database (CMDB), and multiple preset data processing layers, fused data is obtained, wherein the fused data includes integrated data, feature splicing vectors, and decision coefficients. The framework description information is used to characterize the dependency relationship between the system components in the business application system and the target data sources. Constructing a directed graph based on the framework description information and the fused data in the CMDB, the directed graph including nodes and edges connecting the nodes, the nodes representing the target data sources and system components, and the lengths of the edges representing the degree of dependency between the connected nodes; According to the directed graph, the original data and the decision coefficient, an abnormal distribution characteristic parameter of the business application system is obtained through an abnormal weighted propagation model.

2. The method according to claim 1, characterized in that Multiple data processing layers include data integration layer, feature splicing layer and decision processing layer; The fused data is obtained based on the raw data collected from multiple target data sources, the framework description information of the business application system in the CMDB, and the preset multiple data processing layers, including: The data integration layer integrates the original data associated with the same primary key field according to the primary key field in the framework description information to obtain the integrated data, where the primary key field is used to represent the system component; splicing the feature vectors extracted from the original data according to the dependency relationship of the system components represented by the framework description information through the feature splicing layer to obtain the feature splicing vector; The decision processing layer generates the decision coefficient according to the influence of the target data source and the framework description information on the anomaly analysis, and the decision coefficient is used to characterize the importance of the target data source and system components and the uncertainty of the anomaly analysis prediction; The fused data is obtained according to the integrated data, the feature splicing vector and the decision coefficient.

3. The method according to claim 1, characterized in that The step of constructing a directed graph based on the framework description information and the fusion data in the CMDB includes: Generate a dependency matrix based on the framework description information, wherein the dependency matrix represents dependency relationships between functional modules and functional resources of the business application system, wherein the functional modules include some system components, and the functional resources include some system components and the target data source; Generate nodes and edges according to the dependency matrix, where two nodes connected by an edge have a dependency relationship; Based on the fused data, obtaining the length of the edge; The directed graph is constructed according to the nodes, edges and edge lengths.

4. The method according to claim 1, wherein After constructing the directed graph based on the framework description information and the fusion data in the CMDB, the method further includes: Calculate the shortest distance between two nodes connected by the edge according to the length of the edge in the directed graph; For any node, using the shortest distance between the any node and other nodes connected by edges in the directed graph and the number of nodes in the directed graph, obtaining a centrality parameter of the any node; Selecting nodes whose centrality parameters meet preset conditions as key nodes in anomaly analysis; The method of performing multiple iterative operations based on the directed graph, the original data, and the decision coefficient to obtain abnormal distribution characteristic parameters of the business application system includes: Multiple iterative operations are performed based on the key nodes, the original data and the decision coefficients in the directed graph to obtain abnormal distribution characteristic parameters of the business application system.

5. The method according to claim 1, wherein The method of obtaining the abnormal distribution characteristic parameters of the business application system by using an abnormal weighted propagation model based on the directed graph, the original data, and the decision coefficient includes: For any node in the directed graph, performing iterative calculations through the weighted propagation model according to the abnormal parameter of the arbitrary node, the propagation weight of the upstream node of the arbitrary node, the abnormal parameter of the upstream node, and a preset attenuation factor, and updating the abnormal parameter of the arbitrary node until the updated abnormal parameter of the arbitrary node meets a preset convergence condition, wherein the abnormal parameter is obtained based on the original data and the decision coefficient; The abnormal distribution characteristic parameter is calculated based on the updated abnormal parameters of at least some nodes that meet the preset convergence condition. The abnormal distribution characteristic parameter is used to characterize the abnormal concentration degree of the nodes in the directed graph.

6. The method according to claim 5, characterized in that The abnormal distribution characteristic parameter is calculated based on the updated abnormal parameters of at least some nodes that meet the preset convergence condition, including: Normalizing the updated abnormal parameters of at least some of the nodes that meet the preset convergence condition to obtain processed abnormal parameters; Determining the sum of the processed abnormal parameters of at least some of the nodes as a system global abnormal parameter; The abnormal distribution characteristic parameters are obtained according to the system global abnormal parameters and the processed abnormal parameters of the focus nodes, wherein the focus nodes include nodes whose abnormality degrees represented by the processed abnormal parameters in at least some of the nodes meet the preset abnormality conditions.

7. The method according to claim 5, characterized in that After obtaining the abnormal distribution characteristic parameters of the business application system through an abnormal weighted propagation model based on the directed graph, the original data and the decision coefficient, the method further includes: When the abnormal distribution characteristic parameter is higher than or equal to the preset distribution threshold, the system component corresponding to the focus node is subjected to abnormal control processing, and the focus node includes the node whose abnormal degree represented by the updated abnormal parameter meets the preset abnormal condition; When the abnormal distribution characteristic parameter is lower than the preset distribution threshold, the business application system as a whole is subjected to robustness improvement processing.

8. The method according to claim 1, characterized in that Before obtaining fused data based on the raw data collected from the multiple target data sources, the framework description information of the business application system in the CMDB, and the preset multiple data processing layers, at least one of the following is further included: Determine the target data source for the business application system based on user application information, the CMDB's business chain, and compliance screening conditions, and collect raw data from the target data source. The user application information includes user permission information and business application system information. The raw data is preprocessed to obtain the preprocessed raw data, wherein the data preprocessing includes at least one of data cleaning processing, data standardization processing, and semantic association processing, and the semantic association processing is used to establish logical relationships and contextual associations between the raw data.

9. The method according to claim 1, characterized in that After obtaining fused data based on the raw data collected from multiple target data sources, the framework description information of the business application system in the configuration management database CMDB, and multiple preset data processing layers, the method further includes: The historical data in the fused data is processed using an offline calculation method to obtain historical data quality assessment parameters, and the real-time data in the fused data is processed using a real-time calculation method to obtain real-time data quality assessment parameters. Based on the historical data quality assessment parameters and / or the real-time data quality assessment parameters, the fused data that meets the preset quality requirements is screened.

10. The method according to claim 1, characterized in that After constructing the directed graph based on the framework description information and the fusion data in the CMDB, the method further includes: Tracking and detecting dependency relationships between nodes represented by the directed graph to identify dependency anomalies; An exception reduction strategy matching the dependency exception event is executed.

11. The method according to claim 1, wherein Also includes: Determining the current framework structure of the business application system based on actual operation data of the business application system during operation; The framework description information of the CMDB is updated according to the comparison result between the current framework structure and the framework description information of the CMDB.

12. A CMDB-based anomaly analysis device, characterized in that: include: A data fusion module is configured to obtain fused data based on the raw data collected from multiple target data sources, the framework description information of the business application system in the configuration management database (CMDB), and multiple preset data processing layers. The fused data includes integrated data, feature splicing vectors, and decision coefficients. The framework description information is used to characterize the dependency relationships between system components in the business application system and the target data sources. A directed graph construction module is used to construct a directed graph based on the framework description information and the fused data in the CMDB. The directed graph includes nodes and edges connecting the nodes. The nodes are used to represent the target data sources and system components. The length of the edges is used to represent the degree of dependency between the connected nodes. The anomaly analysis module is used to obtain the anomaly distribution characteristic parameters of the business application system through an anomaly weighted propagation model based on the directed graph, the original data and the decision coefficient.

13. A CMDB-based anomaly analysis device, characterized in that: include: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the CMDB-based anomaly analysis method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the CMDB-based anomaly analysis method according to any one of claims 1 to 11 is implemented.

15. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the CMDB-based anomaly analysis method according to any one of claims 1 to 11.