Root cause analysis method and device of database, computer equipment and storage medium

By automatically analyzing cloud platform alarm events and generating database topology and detection portraits, the problem of inefficient cloud database fault location in traditional methods is solved, rapid and accurate root cause analysis is achieved, and the stability and reliability of the cloud platform are improved.

CN120743601APending Publication Date: 2025-10-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510852782.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional cloud database fault location and analysis methods are inefficient and error-prone, making it difficult to quickly and accurately locate the root cause. This impacts the stability and availability of cloud applications, especially when the cloud platform architecture is complex.

Method used

By receiving and parsing alarm events from the cloud platform, generating database topology and association relationships, building detection portraits, and conducting multi-dimensional root cause analysis, the root cause of the fault can be automatically identified.

Benefits of technology

It shortens the fault discovery time, improves the efficiency and accuracy of fault location, reduces misjudgments and missed judgments, and enhances the stability and reliability of the cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743601A_ABST
    Figure CN120743601A_ABST
Patent Text Reader

Abstract

The invention discloses a root cause analysis method and device of a database, computer equipment and a storage medium, and the method comprises the steps: receiving a cloud platform alarm event, and analyzing the alarm event to obtain an analysis result; loading the database architecture from the cloud platform according to the analysis result to generate a database topological structure; associating the database topological structure to generate an association relationship of the storage features; combining the database topological structure and the incidence relation to generate a detection portrait; and performing root cause analysis on the detection portrait to obtain an analysis result. By implementing the method, possibly affected database resources can be quickly positioned, the root cause of the fault can be accurately identified, the requirements of cloud operation and maintenance on the accuracy and timeliness of cloud database fault positioning can be effectively met, the overall efficiency and accuracy of fault positioning are improved, and the fault positioning efficiency is improved. The technical scheme can be applied to the fields of finance and medical health.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database root cause analysis, and more specifically to a database root cause analysis method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the rapid development of information technology in key sectors such as finance and healthcare, cloud computing, as an emerging computing model, is gradually penetrating the information development of various industries. By providing computing, storage, and network resources to users as services, cloud computing technology has significantly reduced enterprise IT costs, improved resource utilization, and promoted the rapid development of business innovation. In recent years, the application of cloud computing has become increasingly widespread, becoming a key infrastructure supporting enterprise digital transformation and internet services.

[0003] However, with the increasing application of cloud computing technology and the continuous expansion of cloud platforms, the architecture and deployment environment of cloud platforms are becoming increasingly complex. Cloud platforms typically include a large number of physical servers, virtual machines, storage devices, network equipment, and various software and services. The interactions and dependencies between these components are complex, and failures in any link can affect the stable operation of the entire cloud platform.

[0004] Especially in the cloud database sector, databases are key components for storing and managing data within cloud platforms. Their stability and reliability are directly linked to the proper functioning of both user and cloud applications. Cloud databases must not only process massive data requests but also ensure data consistency, integrity, and availability. However, due to the complexity and dynamic nature of cloud platform architectures, locating and analyzing cloud database faults is extremely difficult. A cloud database failure can not only render user applications inaccessible but can also impact the availability of the entire cloud application, resulting in significant losses for the enterprise.

[0005] Traditional fault location and analysis methods often rely on human experience and manual troubleshooting. This approach is not only inefficient but also prone to errors, making it difficult to meet the accuracy and timeliness requirements of cloud platforms for fault location and analysis. This is especially true in the case of cloud database failures, as the fault symptoms may involve multiple layers and components, making it difficult for traditional methods to quickly and accurately locate the root cause.

[0006] Therefore, in view of the challenges of locating and analyzing cloud platform faults, especially the difficulties in locating and analyzing cloud database faults, an automated and efficient root cause analysis method is needed to ensure the stable operation of user applications and cloud applications, and thus promote the further development and application of cloud computing technology. Summary of the Invention

[0007] The purpose of the present invention is to overcome the defects of the prior art and provide a root cause analysis method, device, computer equipment and storage medium for a database.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] Database root cause analysis methods, including:

[0010] Receive cloud platform alarm events and analyze them to obtain analysis results;

[0011] Load the database schema from the cloud platform based on the parsing results to generate the database topology;

[0012] Associating the database topology to generate association relationships for storing features;

[0013] Combine database topology and relationships to generate detection profiles;

[0014] Perform root cause analysis on the detection profile to obtain analysis results.

[0015] The present invention also provides a root cause analysis device for a database, comprising:

[0016] The receiving and parsing unit is used to receive the cloud platform alarm events and parse the alarm events to obtain the parsing results;

[0017] A loading unit, used to load the database architecture from the cloud platform according to the parsing results to generate a database topology structure;

[0018] An association unit, used to associate the database topology structure to generate an association relationship for storing features;

[0019] A combining unit, used to combine the database topology and association relationships to generate a detection profile;

[0020] The analysis unit is used to perform root cause analysis on the detection profile to obtain analysis results.

[0021] The present invention further provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0022] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows: by automatically receiving and parsing alarm events from the cloud platform, the database resources that may be affected can be quickly located, effectively shortening the fault discovery time; and combining with the database architecture information loaded from the cloud platform, the database topology structure is automatically generated, further accelerating the definition of the fault scope and improving the overall efficiency of fault location; in addition, by further associating the characteristics of the database and its associated resources on the basis of generating the database topology structure, constructing the association relationship of the storage characteristics, and then combining the database topology structure and the association relationship to generate a comprehensive detection portrait, and conducting in-depth analysis of abnormal phenomena through multi-dimensional detection data, accurately identifying the root cause of the fault, effectively avoiding misjudgment and missed judgment, and enhancing the accuracy of fault analysis.

[0024] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 A schematic diagram of an application scenario of the root cause analysis method for a database provided by an embodiment of the present invention;

[0027] Figure 2 A schematic diagram of a process for root cause analysis of a database provided by an embodiment of the present invention;

[0028] Figure 3 A schematic block diagram of a root cause analysis device for a database provided by an embodiment of the present invention;

[0029] Figure 4 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0031] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0032] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0033] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0034] See also Figure 1 and Figure 2 , Figure 1 A schematic diagram of an application scenario of the root cause analysis method for a database provided by an embodiment of the present invention. Figure 2 A schematic flow chart of a database root cause analysis method provided in an embodiment of the present invention. The database root cause analysis method is applied to a server, which interacts with a terminal for data exchange. By automatically receiving and parsing alarm events from a cloud platform, the server quickly locates potentially affected database resources, effectively shortening the time to discover a fault. Combined with the database architecture information loaded from the cloud platform, the database topology is automatically generated, further accelerating the definition of the fault scope and improving the overall efficiency of fault location. In addition, based on the generation of the database topology, the characteristics of the database and its associated resources are further associated, and an association relationship for storing the characteristics is constructed. Combined with the database topology and the association relationship, a comprehensive detection portrait is generated. The abnormal phenomena are deeply analyzed through multi-dimensional detection data, and the root cause of the fault is accurately identified, effectively avoiding misjudgment and missed judgment, and enhancing the accuracy of fault analysis.

[0035] Figure 2 FIG. 1 is a flow chart of a root cause analysis method for a database provided by an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S150.

[0036] S110, receiving a cloud platform alarm event, and analyzing the alarm event to obtain an analysis result;

[0037] Specifically, alarm events from the cloud platform are received in real time or periodically through the API interface or message queue provided by the cloud platform. The received alarm events are parsed to extract key information, such as alarm resources, alarm indicators, alarm time, etc. Based on the alarm indicators, the category of the alarm resource is determined, such as whether it is a database-related resource (library cluster, library, library instance, etc.). Then, detailed information of the alarm resource is queried from the configuration management database (CMDB), including resource name, location, status, etc. If the alarm resource is not database-related, the associated database is searched from the CMDB based on the resource relevance and placed in the analysis queue. Finally, the alarm resources, alarm indicators, resource categories, detailed information and analysis queues are integrated to obtain the analysis results.

[0038] In other words, by automatically receiving and analyzing alarm events, potentially affected database resources can be quickly located, shortening the time it takes to discover a fault. Furthermore, the automated process reduces the time required for manual analysis and judgment, improving fault location efficiency.

[0039] In one embodiment, receiving a cloud platform alarm event and parsing the alarm event to obtain a parsing result includes:

[0040] Establish an alarm event receiving interface to receive alarm events from the cloud platform in real time or periodically;

[0041] Specifically, based on the communication protocols and interface specifications provided by the cloud platform, select an appropriate interface type, such as a RESTful API or a message queue interface (e.g., RabbitMQ, Kafka, etc.). If the cloud platform supports real-time push notifications, a message queue-based interface is preferred to enable real-time notifications. If the cloud platform only supports periodic notifications, a scheduled task combined with an API interface can be used to periodically retrieve notifications from the cloud platform. Developers should use an appropriate programming language (e.g., Java, Python, etc.) and framework to write the receiving interface code according to the interface specifications. The developed interface should be deployed on a dedicated server or cloud service instance to ensure stable operation, high availability, and scalability to handle the large number of notifications received. To ensure interface security, use authentication mechanisms (e.g., OAuth, API keys, etc.) to authenticate the cloud platform sending notifications, ensuring that only legitimate cloud platforms can send notifications to the interface. Furthermore, encrypt interface communications (e.g., using HTTPS) to prevent theft or tampering of notifications during transmission.

[0042] Analyze the content of received alarm events to extract alarm resources and alarm indicators;

[0043] Specifically, detailed parsing rules are developed based on the format and content characteristics of cloud platform alarm events. For example, if the alarm event is in JSON format, define the meaning of each field and the corresponding parsing logic, such as extracting the alarm resource identifier from the "resource_id" field and extracting the alarm metric name from the "metric_name" field. According to the parsing rules, select the appropriate parsing tool or library. For simple JSON format parsing, you can use the JSON parsing library built into the programming language; for complex custom formats, you may need to develop a dedicated parser. During the parsing process, the content of the alarm event is checked and extracted item by item to ensure that the extracted alarm resources and alarm metrics are accurate. Taking into account that alarm events may have abnormal situations such as format errors and missing data, an exception handling mechanism is added to the parsing process. When encountering an abnormal situation, record detailed error information and mark the alarm event as an exception to facilitate subsequent manual investigation and processing, while ensuring that the parsing process is not interrupted by a single abnormal alarm.

[0044] Based on the alarm indicator, search the resource classification database or configuration file for the corresponding resource category to determine the specific type of alarm resource;

[0045] Specifically, a resource classification database or configuration file is established in advance to divide the various resources potentially involved in the cloud platform according to certain classification standards, and corresponding alarm indicator ranges or characteristics are defined for each resource category. For example, database resources can be divided into categories such as relational databases and non-relational databases, and specific alarm indicators are defined for each category, such as an alarm indicator for the number of connections for relational databases and an alarm indicator for disk space. After receiving the alarm indicator, an appropriate search algorithm (such as a binary search or hash search) is used to quickly locate the corresponding resource category in the resource classification database or configuration file. If the alarm indicator is associated with multiple resource categories, the most appropriate resource category is determined based on preset priority rules or matching algorithms. Considering that cloud platform resources may continue to increase or adjust as business develops, a dynamic update mechanism for the resource classification database or configuration file is established. When new resource types are added to the cloud platform or existing resource categories change, the resource classification database or configuration file is updated promptly to ensure the accuracy of the search results.

[0046] Query the detailed information of the alarm resource from the configuration management database;

[0047] Specifically, a connection is established with the configuration management database (CMDB) through a database connection driver or API interface. During the integration process, necessary authentication and permission configuration are performed to ensure that the system can access the data in the CMDB normally. Based on the identifier of the alarm resource (such as the resource ID), a suitable query statement is constructed to query the detailed information of the resource from the CMDB. The detailed information may include the name of the resource, location, the business system to which it belongs, the person in charge, other associated resources, etc. When constructing the query statement, consider query efficiency and data accuracy to avoid using overly complex queries that may cause performance problems. To improve query efficiency, the detailed information of resources queried from the CMDB is cached. At the same time, a cache update mechanism is established to regularly check the changes in resource information in the CMDB. When it is found that the resource information has been updated, the data in the cache is updated in a timely manner to ensure that the queried resource details are the latest.

[0048] Integrate alarm resources, alarm indicators, resource categories and detailed information to obtain analytical results.

[0049] Specifically, define a unified data structure to store the integrated parsing results. This data structure should include fields such as alarm resources, alarm indicators, resource categories, and detailed information. Design appropriate data types and field lengths based on actual requirements. Additionally, write data integration logic code to integrate the alarm resources and alarm indicators parsed from alarm events, the resource categories retrieved from the resource classification database or configuration files, and the detailed resource information retrieved from the CMDB, all according to the defined data structure. During the integration process, perform necessary data format conversion and validation to ensure that the integrated data meets requirements. Store the integrated parsing results in a designated data storage medium (such as a database, file system, etc.) for subsequent analysis and processing. Furthermore, based on business requirements, output the parsing results to the appropriate system or interface for review and use by operations and maintenance personnel. For example, the parsing results can be pushed to the alarm display interface of the monitoring system or sent to the operation and maintenance personnel's message notification channel.

[0050] In other words, by establishing a dedicated alarm event receiving interface, real-time or periodic alarm event reception is achieved, eliminating the tedious process of manual alarm collection and significantly reducing alarm response time. Furthermore, automated content parsing and resource category search processes quickly and accurately extract key information, reducing manual analysis and judgment time. This allows operations and maintenance personnel to more quickly obtain alarm details, enabling them to take timely action and improving overall alarm processing efficiency. Furthermore, detailed alarm content parsing accurately extracts alarm resources and indicators, preventing inaccurate alarm information caused by manual interpretation errors. Searching for resource categories in the resource classification database or configuration files further clarifies the specific type of alarm resource, helping operations and maintenance personnel more accurately understand the meaning and impact of the alarm. Furthermore, querying detailed information about alarm resources from the configuration management database provides more comprehensive data support for subsequent troubleshooting and resolution, helping to accurately determine the cause of the fault and improve the accuracy of troubleshooting. Furthermore, unified parsing rules and resource classification standards enable standardized parsing and processing of alarm events from diverse sources and formats, achieving standardized management of alarm information. This helps establish a unified alarm monitoring and processing process, making it easier for operations personnel to categorize, count, and analyze alarms. This allows them to better understand the operational status of the cloud platform, promptly identify potential issues and trends, and provide a basis for optimizing and improving the cloud platform. Furthermore, the integrated analysis results provide foundational data for automated operations. By integrating the analysis results with automated operations tools or systems, automated alarm-based troubleshooting processes can be implemented, such as automatically triggering troubleshooting scripts and adjusting resource allocation. This not only improves operations efficiency but also reduces errors caused by human error, enhancing the stability and reliability of the cloud platform.

[0051] S120, loading the database architecture from the cloud platform according to the parsing result to generate a database topology structure;

[0052] Specifically, the database architecture information is first loaded from the cloud platform, including library clusters, libraries, library instances, library roles (master-slave, master-backup, proxy, and other relationships), proxy software, cluster software, and so on. Next, applications that have recently called the library are loaded from the link tracking system. The application architecture and operating status are then loaded from the cloud platform and placed as upstream applications into the database topology. Finally, based on the loaded database architecture and upstream application information, a database topology is generated, which displays the hierarchy and relationships of the database and its associated resources.

[0053] In other words, the generated topology comprehensively displays the hierarchy and relationships of the database and its associated resources, helping operations personnel intuitively understand the database architecture. This topology allows for rapid identification of the potential impact of a failure, providing a foundation for subsequent fault analysis.

[0054] In one embodiment, the step of loading the database architecture from the cloud platform according to the parsing result to generate the database topology structure includes:

[0055] Load the database architecture and operation status information through the API or management interface provided by the cloud platform;

[0056] Specifically, first, identify the API interface documentation provided by the cloud platform for obtaining database architecture and operational status information. According to the documentation, write the API call code in an appropriate programming language (such as Python or Java). When calling the API, authenticate yourself according to the cloud platform's authentication mechanism. Common authentication methods include API keys and OAuth. For example, if using API key authentication, add a field containing the API key to the request header to ensure that the cloud platform can recognize and authorize access requests. Initiate an API request to obtain database architecture information, such as database type (relational, non-relational, etc.), database instance list, table structure, storage engine, etc.; as well as operational status information, such as CPU usage, memory usage, disk I / O, and number of connections. Parse the returned API response data. Based on the data format returned by the API (such as JSON or XML), use the appropriate parsing library or method to extract the required information and store it in an appropriate data structure for subsequent processing. If the cloud platform does not provide a comprehensive API or certain information cannot be obtained through the API, you can use the cloud platform's management interface. Use automated testing tools (such as Selenium) or write browser automation scripts to simulate manual operation to log in to the cloud platform's management interface. In the management interface, locate the database-related page, use page element positioning technology (such as XPath, CSS selector, etc.) to find the elements containing architecture information and operation status information, extract this information and save it.

[0057] Use the link tracking system or the application management function of the cloud platform to identify and load the application information that has called the database in the recent period to obtain related applications;

[0058] Specifically, if the cloud platform has integrated a link tracking system (such as Zipkin, Jaeger, etc.), you first need to obtain access rights and related interface information for the link tracking system. According to the API documentation of the link tracking system, write code to query the call links related to the target database within a specified time range. By setting an appropriate time range and database identifier, filter out call records involving the database. Extract the caller's application information from the call record, such as application name, application instance ID, call frequency, etc., and identify these applications as related applications.

[0059] If the cloud platform provides application management capabilities, access the application management module through the cloud platform's API or management interface. Using the query interface or interface provided by the cloud platform, search for applications that have recently interacted with the database based on database identifiers or associations. For example, by querying the database's connection records or access logs, find the application that initiated the connection or access and identify it as an associated application. Obtain detailed information about the associated application, such as the application architecture (e.g., microservices architecture, monolithic architecture, etc.), application deployment location, and application owner, and store it.

[0060] Load the architecture information and running status of the associated applications from the cloud platform according to the associated applications;

[0061] Specifically, for each associated application, obtain its architecture information based on the API or documentation provided by the cloud platform. For example, if the application is a microservices architecture, obtain the list of microservices, inter-service dependencies, and service registry information; if the application is a monolithic architecture, obtain information such as the application's module division and main functional components. Similarly, the obtained application architecture information is parsed and stored, and can be managed using data structures similar to database architecture information.

[0062] Use the cloud platform's API or monitoring tools to obtain the operating status of associated applications. For example, obtain application metrics such as CPU usage, memory usage, response time, and error rate. For containerized applications provided by some cloud platforms, you can also obtain the container's operating status, such as the number of container instances and container resource allocation. This operating status information is stored in association with the associated application.

[0063] Integrate the database's architecture information and operating status information, as well as the architecture information and operating status of associated applications, to build a database topology.

[0064] Specifically, a unified data structure is designed to represent the database topology. This data structure can take the form of a tree, graph, or relational data table, depending on actual needs and data complexity. The database's architecture and operational status information are added to the data structure as core nodes of the topology. Then, for each associated application, it is added to the topology as a child node associated with the database node, and the associated application's architecture and operational status information is populated into the corresponding child node.

[0065] In the topology, establish relationships between databases and associated applications. For example, use edges to represent application-to-database calls. Edge attributes can include call frequency and call method (e.g., read / write operations). To more intuitively display the database topology, use visualization tools (such as Graphviz and D3.js) to graphically present the integrated data structure. This visualization clearly displays database nodes, associated application nodes, and the relationships between them, making it easier for operations and maintenance personnel to understand and analyze.

[0066] In other words, by loading database architecture and operational status information, operations personnel can gain a detailed understanding of the database's internal structure, configuration, and current operational status. Furthermore, by obtaining information about associated applications, operations personnel can gain a comprehensive understanding of the database's operating environment, including which applications are using the database, their architecture, and their operational status. This helps comprehensively assess database performance and stability, and proactively identify potential issues. Furthermore, when a database failure occurs, the constructed database topology clearly displays the relationship between the database and its associated applications. Operations personnel can quickly identify which applications may be affected by the failure, allowing for targeted troubleshooting and resolution. For example, if a database connection issue occurs, the topology allows personnel to quickly identify applications that depend on the database and check their operational status to determine whether the database failure is causing the application problem, improving troubleshooting efficiency. Furthermore, by integrating and analyzing the operational status information of the database and its associated applications, operations personnel can understand database load and the utilization of database resources by each application. For example, if you find that frequent queries from an application cause excessive database CPU usage, you can optimize the application based on the analysis results, or adjust the database resource allocation strategy, such as increasing the database's CPU resources or optimizing query statements, to improve the database's overall performance and resource utilization.

[0067] S130, associating the database topology structure to generate an association relationship for storing features;

[0068] Specifically, we first extract key features from the database topology, such as library roles, running hosts, and associated switches. Next, we construct a three-dimensional array of associations to store these features. The first dimension represents the association features, and the second and third dimensions represent the association keys and attributes, respectively. Finally, we construct an association based on the degree of association between the features.

[0069] In other words, by building association relationships, the analysis of the database topology can be refined and the accuracy of fault location can be improved.

[0070] In one embodiment, associating the database topology structure to generate an association relationship for storing features includes:

[0071] According to the analysis requirements, determine the features that need to be associated to obtain the feature set;

[0072] Specifically, based on the actual cloud database operation and maintenance scenarios and business objectives, conduct an in-depth analysis of the specific requirements for database topology correlation analysis. For example, if the business is focused on database performance bottlenecks and fault propagation paths, it may be necessary to correlate database performance metrics such as CPU usage, memory usage, disk I / O, and number of connections, as well as topological features such as dependencies between database instances and application-to-database call relationships. Based on the requirements analysis results, select features closely related to the target analysis from a wide range of possible features to form a feature set. During the screening process, consider factors such as feature importance, availability, and quantifiability. For example, for database performance analysis, identify CPU usage, memory usage, and the number of slow queries as key performance features. For fault propagation analysis, identify master-slave relationships between database instances and application-to-database connections as key topological features. Detailed definitions should be provided for these identified features, including their meaning, calculation method, data source, and value range. Document these definitions to ensure a consistent understanding of the features among all project participants, providing clear guidance for subsequent feature extraction and correlation analysis.

[0073] Extracting detailed information of each feature from the database topology structure according to the feature set to obtain feature information data;

[0074] Specifically, select appropriate data access interfaces and tools based on the database topology storage method (e.g., relational database, graph database, etc.) and the feature data source. For example, if the topology is stored in a relational database, SQL statements can be used for data queries; if it is stored in a graph database, graph database query languages ​​(e.g., Cypher, Gremlin, etc.) can be used for feature extraction. For each feature, code the corresponding feature extraction logic. The code should clearly specify how to obtain detailed information about the feature from the database topology. For example, for the CPU usage feature, query the performance monitoring table of the database instance to obtain CPU usage data within a specified time range. For the dependency feature between database instances, query the edge information in the topology to obtain dependency relationships between instances, such as master-slave and read-write splitting. The extracted feature information data should be cleaned and preprocessed to remove noise, handle missing values, and perform data format conversion. For example, missing CPU usage data can be filled using interpolation. In the case of inconsistent data formats, unified conversion is performed to ensure the accuracy of subsequent analysis.

[0075] Based on the feature information data, associations between features are established to generate association relationships for storing features.

[0076] Specifically, combining business knowledge and data analysis experience, association rules between features are determined. Association rules can be developed based on factors such as causality, correlation, and time sequence. For example, a sudden increase in database CPU usage may be associated with an increase in the number of slow queries or a sudden increase in the number of requests from a related application. These can serve as the basis for association rules. Furthermore, appropriate association algorithms are selected based on the association rules and data characteristics. Common association algorithms include rule-based association algorithms (such as the Apriori algorithm) and similarity-based association algorithms. The selected algorithm is used to process feature information data and establish associations between features. For example, the Apriori algorithm can be used to mine frequent itemsets, identifying frequently occurring feature combinations and thus determining associations between features. Established associations are stored in appropriate data structures, such as association matrices and association graphs. Furthermore, to facilitate subsequent analysis and querying, associations can be stored in database tables containing information such as associated feature pairs, association strength, and association type. Furthermore, visualization tools can be used to graphically display associations, such as using network diagrams to visualize the association network between features, intuitively presenting the complex relationships between features.

[0077] In other words, by establishing correlations between features, the causes of database failures can be analyzed more comprehensively. For example, when a database has performance issues, it can be analyzed not only from a single performance indicator (such as CPU usage), but also by combining related features (such as the number of slow queries, the number of requests from related applications, etc.) for a comprehensive judgment, thereby more accurately locating the root cause of the failure and improving the accuracy of fault diagnosis. In addition, understanding the correlations between features helps to discover database performance bottlenecks and potential performance optimization points. For example, through analysis, it is found that frequent queries from a certain related application cause the database CPU usage to be too high. The application can be optimized, such as optimizing query statements, increasing cache, etc., thereby reducing the load on the database and improving overall performance.

[0078] S140, combining the database topology and association relationships to generate a detection profile;

[0079] Specifically, the system first performs inspections across multiple dimensions, including the database itself, associated resources, and upstream applications, collecting data on changes, maintenance, alarm events, status, and performance. Next, it aggregates this multi-dimensional inspection data to form an inspection profile of the database and its associated resources. Finally, based on this aggregated data, it generates a detailed inspection profile that displays the current status of the database and its associated resources.

[0080] In other words, the detection portrait comprehensively displays the current status of the database and its associated resources, helping operation and maintenance personnel to quickly understand the overall situation, providing a rich data foundation for subsequent root cause analysis, and helping to improve the accuracy and comprehensiveness of the analysis.

[0081] In one embodiment, combining the database topology and association relationships to generate a detection profile includes:

[0082] Collect the database's latest change records, maintenance operations, alarm events, service status, performance indicators, and number of connections from the database topology to obtain self-collected information;

[0083] Specifically, access the cloud platform's change management system or the database's own change logging module. By writing scripts or using the platform's API, query database change records by time range (e.g., the past week, month, etc.), including information such as database schema adjustments (e.g., table structure changes, index additions and deletions), and configuration parameter modifications (e.g., memory allocation, connection limits), and store these records in a temporary data structure. Obtain database maintenance operation records from the cloud platform's operation and maintenance management system, including information such as the time, performer, and results of database backups, restores, and patch installations. Similarly, organize this information and store it in a temporary data structure. Subscribe to the cloud platform's alarm system to receive database-related alarm events in real time or periodically. Parse the alarm events, extract key information such as the alarm type (e.g., performance alarm, availability alarm), alarm time, and alarm level, and add them to the temporary data structure. Use the cloud platform's monitoring tools or the database's built-in monitoring interface to regularly query the database's service status, such as the database instance's operating status (running, stopped, faulty, etc.) and service availability metrics (e.g., service response time, success rate, etc.), and record this status information. Use the database's performance monitoring module or the cloud platform's performance monitoring service to obtain key database performance indicators, such as CPU usage, memory usage, disk I / O, query response time, and throughput. Collect these metrics at regular intervals (e.g., every minute, every hour) and store them in a temporary data structure. Use the database's management interface or monitoring tools to obtain real-time database connection information, including the current number of connections, maximum number of connections, and connection timeouts, and add this information to your collected information.

[0084] Collect relevant change records, maintenance operations, alarm events, status, and performance information from resources associated with the database based on the association relationships to obtain associated collection information;

[0085] Specifically, identify the resources directly associated with the database, such as storage devices, network devices, server hosts, etc. For each associated resource, access its corresponding change management system or log records, and query change records within the same time range as the time range of the database's own information collection. For example, the expansion of storage devices, configuration changes of network devices, etc., extract these records. Obtain maintenance operation records from the operation and maintenance management system of associated resources, such as inspections of storage devices, upgrades and maintenance of network devices, and other operational information, and organize and store them. Subscribe to the alarm system of associated resources, collect alarm events generated by associated resources within the time period associated with the database, and analyze the possible impact of alarm events on the database, such as storage device failure alarms that may cause abnormal reading and writing of database data. Use corresponding monitoring tools or interfaces to collect status and performance information of associated resources. For example, collect indicators such as the remaining space and read and write speed of storage devices, and the bandwidth utilization and latency of network devices, and use this information as part of the associated collected information.

[0086] Use application performance management tools or logs to collect change records, release information, and database call performance information of upstream applications associated with the database to obtain application collection information.

[0087] Specifically, based on the database topology and association relationships, determine the upstream applications that have a calling relationship with the database. Access the change management system used by the application development team or operation and maintenance team to query the change records of the upstream application within a specified time range, such as code modifications, function updates, configuration adjustments, and other information, with a focus on changes related to database interaction. Obtain the application's release information from the application's version management system, including release time, release content (such as new features, fixed database-related issues, etc.), release manager, etc., and organize and collect this information. Use application performance management tools (such as APM tools) or database call logs to collect performance information on upstream applications calling the database, such as call frequency, average response time, call success rate, slow query calls, etc. Analyze and organize these data to obtain application collection information.

[0088] Integrate self-collected information, related collection information and application collection information to generate a detection profile.

[0089] Specifically, a unified data structure is designed to store the integrated detection profile information. This data structure can adopt a hierarchical structure, such as using the database as the core node, storing self-collected information, associated collected information, and application collected information as sub-nodes, and each sub-node is further subdivided into different information categories (such as change records, alarm events, etc.). Data integration code is written to integrate the self-collected information, associated collected information, and application collected information according to the designed data structure. During the integration process, data is deduplicated and formatted to ensure accuracy and consistency. For example, for alarm events occurring at the same time, duplicate records are avoided during integration. The integrated data is stored in a dedicated database or file system to generate a detection profile. At the same time, to facilitate subsequent query and analysis, the detection profile can be indexed and categorized. In addition, visualization tools can be used to display the detection profile in a graphical format, such as using a dashboard to display key indicators and a timeline to display change records and alarm events, so that operation and maintenance personnel can more intuitively understand the overall operation status of the database.

[0090] In other words, by collecting comprehensive information about the database itself, associated resources, and upstream applications, the generated detection profile comprehensively reflects the database's operational status. Operations and maintenance personnel can use the detection profile to obtain detailed information such as database performance metrics, service status, change history, and the impact of associated resources and upstream applications on the database. This provides a deeper understanding of the database's overall operational status and allows them to promptly identify potential issues and risks. Furthermore, when a database failure occurs, the rich information integrated into the detection profile provides powerful support for fault location. Based on the alarm events, change history, call performance, and other information in the detection profile, operations and maintenance personnel can quickly narrow down the scope of the failure and locate the possible cause. For example, if the detection profile shows that the database experienced performance degradation at a certain point in time, and at the same time, a storage device in the associated resources reported an abnormal alarm, and upstream applications had a large number of slow query calls to the database, operations and maintenance personnel can quickly determine that the failure is likely related to the storage device or upstream application calls, greatly improving fault location efficiency.

[0091] S150: Perform root cause analysis on the detection profile to obtain analysis results.

[0092] Specifically, we first identify surface issues from the detection profile, such as failed library instance logins or fully loaded library instances. We then analyze the correlation between the surface issues and the database, associated resources, and upstream applications. Based on this correlation, we add weighted edges between the detection results. Finally, by analyzing the graph formed by these weighted edges, we can trace the root cause from the surface issues and their weights to determine the underlying cause.

[0093] In other words, root cause analysis can accurately pinpoint the root cause of a fault, improving the efficiency and accuracy of troubleshooting. Furthermore, root cause analysis results provide cloud operations personnel with clear troubleshooting guidelines, helping to optimize operations processes and improve the stability and availability of the cloud platform.

[0094] In one embodiment, performing root cause analysis on the detection profile to obtain analysis results includes:

[0095] Test the database itself, database-related resources, and upstream applications based on the test profile to obtain test results for the database itself, database-related resources, and upstream applications.

[0096] Specifically, database self-detection: extract database performance indicator data from the detection portrait, such as CPU usage, memory usage, disk I / O, query response time, etc. Use the preset performance threshold to detect each indicator. If an indicator exceeds the threshold, it is marked as abnormal. For example, if the CPU usage exceeds 80% continuously and lasts for more than 5 minutes, it is determined that the CPU usage is abnormal. Check the operating status of the database instance, such as whether it is running normally, whether there is a failure or restart, etc. By interacting with the cloud platform's monitoring system or the database's own status interface, obtain real-time status information, and compare it with the expected status to determine whether the service status is normal. Analyze the database change records in the detection portrait, check whether the changes comply with the specifications, and whether there are any change operations that may cause failures. For example, check whether the table structure changes have been fully tested, whether the configuration parameter modifications are within a reasonable range, etc.

[0097] Database-related resource detection: For storage devices associated with the database, detect indicators such as remaining space, read and write performance, and disk health status. Obtain relevant data through the storage device's management interface or monitoring tool, and compare it with the preset threshold to determine whether there are any abnormalities in the storage device. Check network equipment's bandwidth utilization, latency, packet loss rate and other indicators to ensure a stable network connection. Use network monitoring tools to monitor network devices in real time, analyze network traffic and performance data, and identify network problems. Obtain the server host's CPU, memory, disk and other resource usage, as well as abnormal information in the system log. Collect data through server management tools or monitoring systems to determine whether the server host has performance bottlenecks or failures.

[0098] Upstream application testing: Analyze the performance of database calls from upstream applications in the test profile, such as call frequency, average response time, and call success rate. Abnormally high call frequency or prolonged response time may indicate performance issues in the upstream application or inappropriate database calls. Review the upstream application's release history to analyze whether the release is relevant to the database and whether any changes have caused database issues. For example, determine whether new database query statements added in the release are efficient and impact database performance.

[0099] Identify the detection portrait to get the apparent problem;

[0100] Specifically, a series of rules are predefined to identify surface problems in detection portraits. For example, it is defined that when multiple performance indicators of the database are abnormal (such as high CPU usage, long query response time) and accompanied by alarm events, it is identified as a "database performance degradation" surface problem; when the number of database connections is detected to fluctuate frequently and exceeds the normal range, it is identified as a "database connection abnormality" surface problem. If the amount of data is large and the rules are difficult to cover all situations, a machine learning model can be used to train and identify detection portraits. Collect historical detection portrait data, mark the surface problems, and use them as a training set. Use a suitable machine learning algorithm (such as decision tree, neural network, etc.) to train the model, and then use the trained model to identify surface problems in new detection portraits.

[0101] Perform correlation analysis between the appearance problem and the database's own detection results to obtain the database correlation degree, and add a weighted edge between the appearance problem and the database's own detection results based on the database correlation degree, namely the database weight;

[0102] Specifically, analyze the degree of correlation between the apparent problem and the database's own detection results. For example, if the apparent problem is "database performance degradation," and the database's own detection results show abnormally high CPU usage and long query response times, calculate the correlation between the two. The correlation can be calculated based on factors such as the number of abnormal indicators, the degree of abnormality, and the temporal consistency of the abnormality. For example, the greater the number of abnormal indicators, the higher the degree of abnormality, and the stronger the temporal consistency, the higher the correlation. Based on the correlation calculation results, determine the weight of the weighted edge. The weight can be expressed as a numerical value, such as the higher the correlation, the greater the weight. For example, if the database correlation is 0.8, the corresponding database weight can be set to 0.8.

[0103] Perform correlation analysis on the appearance problem and the database-related resource detection result to obtain the database-related resource correlation degree, and add a weighted edge between the appearance problem and the database-related resource detection result according to the database-related resource correlation degree, namely, the database-related resource weight;

[0104] Specifically, we similarly analyze the correlation between the perceived issue and the database-related resource detection results. For example, for the perceived issue of "degraded database performance," if the associated resource has insufficient free space on the storage device and reduced read / write performance, we calculate the correlation between the two. This correlation calculation method is similar to the database correlation calculation, taking into account factors such as the impact of resource anomalies on database performance. For example, if the database-related resource correlation is 0.6, the corresponding database-related resource weight is set to 0.6.

[0105] Perform correlation analysis between the appearance problem and the upstream application detection results to obtain the upstream application correlation degree. Based on the upstream application correlation degree, add a weighted edge between the appearance problem and the upstream application detection results, namely the upstream application weight.

[0106] Specifically, the correlation between the apparent problem and the upstream application detection results is analyzed. Factors such as the upstream application's call performance and release changes are analyzed to determine their relationship to the apparent problem. For example, if the apparent problem is an abnormal database connection, and the upstream application detection results show an abnormally high call frequency, the correlation between the two is calculated. For example, if the upstream application correlation is 0.7, the corresponding upstream application weight is set to 0.7.

[0107] Integrate the surface issues, database test results, database-related resource test results, upstream application test results, database weights, database-related resource weights, and upstream application weights to construct a weight graph;

[0108] Specifically, a graph structure is designed to represent the weighted graph, consisting of nodes and edges. Nodes represent the problem representation, the database's own test results, the test results of database-related resources, and the test results of upstream applications. Edges represent the relationships between them, and the edge weights are the weights of each component calculated previously. Using graph theory algorithms and data structures, such as adjacency lists or adjacency matrices, the problem representation, the test results of each component, and the weighted edges are integrated into the graph structure to construct the weighted graph.

[0109] The weight graph is analyzed to obtain analysis results.

[0110] Specifically, a graph traversal algorithm (such as depth-first search, breadth-first search) is used to analyze the weight graph. Starting from the apparent problem node, traverse the edges and nodes connected to it, and find the most likely root cause of the fault based on the weight of the edge and the information of the node. For example, during the traversal process, nodes connected by edges with larger weights are preferentially selected as possible root cause nodes. Analyze the path from the apparent problem to the potential root cause node in the weight graph, and find the path with the largest weight. The information represented by the nodes and edges on this path is more likely to be the root cause of the apparent problem. For example, if starting from the apparent problem "database performance degradation", there is a path passing through the two nodes of "abnormally high database CPU usage" and "upstream applications frequently initiate complex queries", and the weight of this path is the largest, then the analysis results may indicate that the upstream application frequently initiates complex queries, resulting in excessively high database CPU usage, which in turn causes database performance degradation.

[0111] In other words, by comprehensively considering the test results of the database itself, related resources, and upstream applications, as well as their correlation with the underlying problem, a weighted graph is constructed for analysis. This allows for a more comprehensive consideration of the impact of various factors on database failures, avoiding the limitations of single-factor analysis and thus improving the accuracy of root cause analysis. Furthermore, the weighted graph's analysis algorithm can quickly traverse the graph structure to identify the most likely root cause nodes and critical paths. Based on the analysis results, operations and maintenance personnel can directly locate the root cause of the failure, significantly reducing troubleshooting time and improving operations and maintenance efficiency. Furthermore, for complex failures involving the interaction of multiple factors, the weighted graph clearly displays the relationships and weights between each factor, helping operations and maintenance personnel understand the failure's propagation path and impact range. By analyzing the weighted graph, key links in complex failures can be identified, providing a basis for developing effective solutions.

[0112] In one embodiment, after performing root cause analysis on the detection profile to obtain analysis results, the method further includes:

[0113] Visualize the analysis results.

[0114] Specifically, based on the root cause analysis results, determine the key information to display. For example, display the fault's symptoms, possible root cause nodes, the relationships between nodes, and the correlation weights. Also, consider whether to display historical data comparisons and trend analysis to help operations personnel better understand the fault's development process. Furthermore, choose an appropriate chart type based on the content being presented. To display nodes and relationships, use a network graph. Nodes include the symptoms, database test results, database-related resource test results, and upstream application test results. Weighted edges represent the relationships between them. The thickness or color of the edges can indicate the weights. To display data such as performance metrics and correlation values, use bar charts, line charts, and pie charts. For example, use a bar chart to compare the correlations between different root cause nodes, and a line chart to display the trends of performance metrics over time. Furthermore, preprocess the root cause analysis data and convert it into a format that can be recognized by visualization tools or platforms. For example, for network diagrams, node and edge information needs to be organized into a specific JSON format; for bar charts and line charts, the data needs to be organized into a table format containing x-axis and y-axis data. Bind the preprocessed data to the visualization chart. In general visualization libraries, data is associated with chart elements by writing code; in professional operation and maintenance visualization platforms, data sources and data fields can usually be selected through the configuration interface to achieve automatic data binding. Call the rendering function of the visualization tool or platform to generate the visualization chart. During the rendering process, adjust and optimize the chart according to the designed layout and style to ensure that the final visualization effect meets the expectations.

[0115] In other words, through visualization, complex root cause analysis results are presented in intuitive graphs and charts. This allows operations personnel to quickly understand the fault's symptoms, possible root causes, and the relationships between various factors without having to deeply analyze large amounts of data and text. This significantly reduces the difficulty of understanding information and improves operations personnel's work efficiency. Furthermore, the visual interface provides a clear display of information, allowing operations personnel to quickly grasp key fault information and make decisions quickly. For example, by viewing the edges and nodes with higher weights in the network diagram, operations personnel can quickly identify the most likely root cause of the fault and implement appropriate remedial measures, thereby reducing the duration of the fault's impact on business.

[0116] For example, in the financial sector, a bank's core transaction system is crucial for business operations. The stability and performance of its database are directly related to customer fund security and transaction experience. Imagine a bank's core transaction system database experiences slow transaction response times, causing some customer transactions to fail and disrupting normal business operations. In this case, the aforementioned database root cause analysis method is needed to quickly locate and resolve the issue.

[0117] The bank's operations team established a dedicated alarm event receiving interface on the cloud platform. This interface is integrated with the bank's core transaction system monitoring system and can receive alarm events from the cloud platform in real time. Upon receiving an alarm event, the system automatically parses the alarm content. For example, it determines that the alarm resource is a database instance of the core transaction system, and the alarm indicator is that the transaction response time exceeds a preset threshold (for example, 2 seconds). Based on the alarm indicator "transaction response time," the corresponding resource category is searched in the resource classification database and determined that the alarm resource is a database resource, specifically a relational database. The bank's configuration management database (CMDB) is then retrieved for detailed information about the database instance, including the database version, deployment location, storage capacity, and associated application systems. The system integrates the alarm resource (database instance), alarm indicator (transaction response time exceeding the threshold), resource category (relational database), and detailed information (version, deployment location, etc.) to generate the parsed result.

[0118] Through the API provided by the cloud platform, the architecture information of the core transaction system database, such as database table structure, index information, stored procedures, etc., as well as operational status information such as CPU usage, memory usage, and disk I / O, are loaded. Using the link tracking system, the application information that has recently called the database is identified. It is found that in addition to the core transaction system itself, the report generation system and risk assessment system also call the database. Based on the identified related applications, the architecture information and operational status of the report generation system and risk assessment system are loaded from the cloud platform, such as the application deployment architecture, number of service instances, response time, etc. The architecture information and operational status of the database, as well as the architecture information and operational status of the related applications, are integrated to construct a database topology. This topology clearly shows the relationship between the database and related applications, as well as their respective operational status.

[0119] Based on the analysis requirements, determine the features that need to be associated, such as the database's CPU usage, memory usage, disk I / O, transaction response time, and the call frequency and concurrency of associated applications, to form a feature set. Based on the feature set, extract detailed information for each feature from the database topology. For example, extract the database's CPU usage data, memory usage data, and call frequency data of associated applications before and after a failure. Based on the feature information data, establish associations between features. For example, if it is found that when the call frequency of the report generation system suddenly increases, the database's CPU usage will also increase. Thus, an association relationship is established between them, generating an association relationship for storing features.

[0120] Collect recent database change records (such as modifications to database table structures), maintenance operations (such as database backups), alarm events (such as previous transaction response time alarms), service status (such as the operating status of the database instance), performance indicators (such as CPU usage, memory usage, and disk I / O), and connection count information from the database topology to obtain self-collected information. Based on the association relationships, collect related change records (such as storage device expansion), maintenance operations (such as network device inspections), alarm events (such as storage device insufficient space alarms), status and performance information (such as storage device read and write speeds and network device bandwidth utilization) from resources associated with the database (such as storage devices and network devices) to obtain associated collection information. Through application performance management tools or logs, collect change records (such as application version updates), release information (such as the launch of new functions), and database call performance information (such as call frequency and response time) of upstream applications associated with the database (report generation systems and risk assessment systems) to obtain application collection information. Integrate self-collected information, associated collection information, and application collection information to generate a detection profile. This detection portrait comprehensively displays the status and performance information of the database, its associated resources, and upstream applications.

[0121] Based on the detection profile, the database itself, its associated resources, and upstream applications are tested. The detection profile is identified, and the apparent problem is "Excessive transaction response time in the core transaction system database." Correlation analysis is performed between the apparent problem and the database's own detection results, resulting in a database correlation of 0.8. A weighted edge is added between the apparent problem and the database's own detection results, assigning a database weight of 0.8. Correlation analysis is performed between the apparent problem and the database's associated resource detection results, resulting in a database correlation of 0.6. A weighted edge is added between the apparent problem and the database's associated resource detection results, assigning a database weight of 0.6. Correlation analysis is performed between the apparent problem and the upstream application detection results, resulting in an upstream application correlation of 0.7. A weighted edge is added between the apparent problem and the upstream application detection results, assigning a upstream application weight of 0.7. The apparent problem, the database's own detection results, the database's associated resource detection results, the upstream application detection results, the database's weight, the database's associated resource weight, and the upstream application weight are integrated to form a weighted graph. Through the analysis of the weight graph, it was found that the high database CPU usage and the increased frequency of report generation system calls were the key factors leading to long transaction response time.

[0122] The bank's operations team chose Grafana as a visualization tool because it integrates well with the bank's monitoring system and allows for easy access to and display of analysis results. A visualization interface was designed, using a network diagram to illustrate the relationships between surface issues, database test results, test results for database-related resources, and test results for upstream applications. The thickness of the edges represents the weight of the edges. Furthermore, bar charts were used to display the changing trends of performance indicators such as database CPU usage and memory usage, as well as the changing call frequency of the report generation system. The root cause analysis data was preprocessed, converted into a format recognizable by Grafana, and bound to the corresponding chart. Grafana's rendering function was used to generate visualization charts. Node interactivity was implemented, such as hovering the mouse to display node details and clicking a node to expand or collapse related nodes. Furthermore, chart linkage was implemented. When clicking a database CPU usage node in the network diagram, the associated bar chart automatically updates to display the changing trend of CPU usage.

[0123] Based on the root cause analysis, the bank's operations team implemented the following measures: They optimized the database, adjusted database parameters, freed up some memory resources, and reduced CPU usage. Furthermore, they optimized the report generation system, limiting the frequency of its database calls to avoid excessive pressure on the database. Furthermore, they inspected and maintained storage devices to improve their read and write speeds. These measures restored transaction response times in the core transaction system database, resolving the issue and ensuring the normal operation of the bank's core business.

[0124] For example, in the healthcare sector, a hospital's electronic medical record system is a core business system that stores critical patient information, such as medical records, diagnosis results, and medication information. The stability and performance of this system's database are crucial. A failure can prevent doctors from timely accessing patient information, impacting medical decision-making and treatment. Imagine a hospital's electronic medical record system experiences slow data queries, resulting in lengthy wait times for doctors to access patient records, severely impacting medical efficiency. In this case, the aforementioned database root cause analysis method is necessary to quickly identify and resolve the cause of the failure.

[0125] The hospital's information technology department has established a dedicated alarm event receiving interface on the cloud platform. This interface is integrated with the electronic medical record system's monitoring system and can receive alarm events from the cloud platform in real time. Upon receiving an alarm event, the system automatically parses the alarm content. For example, it determines that the alarm resource is a database instance in the electronic medical record system, and the alarm indicator is that the data query response time exceeds a preset threshold (e.g., 3 seconds). Based on the alarm indicator "data query response time," the corresponding resource category is searched in the resource classification database, confirming that the alarm resource is a database resource, specifically a relational database. The hospital's configuration management database (CMDB) is then retrieved for detailed information about the database instance, including the database version, deployment location, storage capacity, and associated application systems (e.g., electronic medical record system, medical imaging system, etc.). The alarm resource (database instance), alarm indicator (data query response time exceeding the threshold), resource category (relational database), and detailed information (version, deployment location, etc.) are integrated to generate the parsed result.

[0126] Through the API provided by the cloud platform, the architectural information of the electronic medical record system database is loaded, such as the database table structure (patient information table, medical record table, etc.), index information, stored procedures, etc., as well as operational status information such as CPU usage, memory usage, disk I / O, etc. Utilizing the link tracking system or the application management function of the cloud platform, the application information that has called the database in the recent period is identified. It is found that in addition to the electronic medical record system itself, the medical imaging system also calls the database to query the patient's basic information. Based on the identified related applications, the architectural information and operational status of the medical imaging system are loaded from the cloud platform, such as the application deployment architecture, number of service instances, response time, etc. The architectural information and operational status of the database, as well as the architectural information and operational status of the related applications, are integrated to construct a database topology. This topology clearly shows the relationship between the database and the related applications, as well as their respective operational status.

[0127] Based on the analysis requirements, the features that need to be associated are determined, such as the database's CPU usage, memory usage, disk I / O, data query response time, and the call frequency and concurrency of associated applications, to form a feature set. Based on the feature set, detailed information for each feature is extracted from the database topology. For example, the database's CPU usage data, memory usage data, and call frequency data for associated applications before and after a failure are extracted. Based on the feature information data, associations are established between the features. For example, if the call frequency of a medical imaging system suddenly increases, the database's CPU usage will also increase. This allows the association between them to be established, generating an association for storing features.

[0128] From the database topology, recent database change records (such as modifications to the database table structure), maintenance operations (such as database backups), alarm events (such as previous data query response time alarms), service status (such as the operating status of the database instance), performance indicators (such as CPU usage, memory usage, and disk I / O), and connection count information are collected to obtain self-collected information. Based on the association relationships, relevant change records (such as storage device expansion), maintenance operations (such as network device inspections), alarm events (such as storage device insufficient space alarms), status and performance information (such as storage device read and write speeds and network device bandwidth utilization) are collected from resources associated with the database (such as storage devices and network devices) to obtain associated collection information. Through application performance management tools or logs, change records (such as application version updates), release information (such as the launch of new features), and database call performance information (such as call frequency and response time) of upstream applications associated with the database (electronic medical record systems and medical imaging systems) are collected to obtain application collection information. Self-collected information, associated collection information, and application collection information are integrated to generate a detection profile. This detection portrait comprehensively displays the status and performance information of the database, its associated resources, and upstream applications.

[0129] Based on the detection profile, the database itself, its associated resources, and upstream applications are tested. The detection profile is identified, and the apparent problem is "slow data query in the electronic medical record system database." Correlation analysis is performed between the apparent problem and the database's own detection results, resulting in a database correlation of 0.75. A weighted edge is added between the apparent problem and the database's own detection results, assigning a database weight of 0.75. Correlation analysis is performed between the apparent problem and the database's associated resource detection results, resulting in a database correlation of 0.6. A weighted edge is added between the apparent problem and the database's associated resource detection results, assigning a database weight of 0.6. Correlation analysis is performed between the apparent problem and the upstream application detection results, resulting in an upstream application correlation of 0.7. A weighted edge is added between the apparent problem and the upstream application detection results, assigning a upstream application weight of 0.7. The apparent problem, the database's own detection results, the database's associated resource detection results, the upstream application detection results, the database weight, the database's associated resource weight, and the upstream application weight are integrated to form a weighted graph. Through the analysis of the weight graph, it was found that the high database CPU usage and the increased call frequency of the medical imaging system were the key factors leading to slow data query.

[0130] The hospital's information department chose to use professional data visualization tools, such as Tableau, which can easily display complex data relationships and analysis results. A visualization interface was designed, using network diagrams to display the relationships between surface issues, database test results, database-related resource test results, and upstream application test results. The thickness of the edges indicates the weight. Line charts were also used to display the changing trends of performance indicators such as database CPU usage and memory usage, as well as the changes in the call frequency of the medical imaging system. The root cause analysis result data was preprocessed, converted into a data format that Tableau could recognize, and bound to the corresponding chart. Tableau's rendering function was called to generate a visual chart. In addition, node interaction functions were implemented, such as hovering the mouse to display node details and clicking a node to expand or collapse associated nodes. At the same time, a chart linkage function was implemented. When the database CPU usage node was clicked in the network diagram, the related line chart was automatically updated to display the changing trend of CPU usage.

[0131] Based on the root cause analysis, the hospital's information technology department implemented the following measures: Optimize the database, adjust database parameters, free up some memory resources, and reduce CPU usage. Optimize the medical imaging system to limit the frequency of database calls to avoid excessive pressure on the database. Inspect and maintain storage devices to improve their read and write speeds. These measures restored data query speeds in the electronic medical record system database, resolving the issue and ensuring the normal operation of the hospital's medical services.

[0132] The above-mentioned database root cause analysis method automatically receives and analyzes alarm events from the cloud platform, quickly locates potentially affected database resources, effectively shortens fault discovery time, and automatically generates a database topology structure based on the database architecture information loaded from the cloud platform, further accelerating the definition of the fault scope and improving the overall efficiency of fault location. In addition, based on the generated database topology structure, the characteristics of the database and its associated resources are further associated, and the association relationship of the storage characteristics is constructed. Combined with the database topology structure and the association relationship, a comprehensive detection portrait is generated. Through multi-dimensional detection data, an in-depth analysis of abnormal phenomena is carried out to accurately identify the root cause of the fault, effectively avoiding misjudgment and missed judgment, and enhancing the accuracy of fault analysis.

[0133] Figure 3 FIG. 3 is a schematic block diagram of a root cause analysis device 300 for a database provided by an embodiment of the present invention. Figure 3As shown, corresponding to the above database root cause analysis method, the present invention also provides a database root cause analysis device 300. The database root cause analysis device 300 includes a unit for executing the above database root cause analysis method, and the device can be configured in a server. Figure 3 , the root cause analysis device 300 of the database includes:

[0134] The receiving and parsing unit 301 is used to receive the cloud platform alarm event and parse the alarm event to obtain the parsing result;

[0135] The loading unit 302 is used to load the database architecture from the cloud platform according to the parsing result to generate a database topology structure;

[0136] The association unit 303 is used to associate the database topology structure to generate an association relationship for storing features;

[0137] A combining unit 304 is used to combine the database topology and the association relationship to generate a detection profile;

[0138] The analysis unit 305 is used to perform root cause analysis on the detection profile to obtain an analysis result.

[0139] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the root cause analysis device 300 and each unit of the above database can refer to the corresponding description in the above method embodiment, and for the convenience and brevity of description, it will not be repeated here.

[0140] The root cause analysis device 300 of the database can be implemented in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer device shown.

[0141] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.

[0142] See Figure 4 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0143] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can cause the processor 502 to perform a database root cause analysis method.

[0144] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0145] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a root cause analysis method of a database.

[0146] The network interface 505 is used to communicate with other devices through the network. Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0147] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:

[0148] Receive cloud platform alarm events and parse them to obtain analysis results; load the database architecture from the cloud platform based on the analysis results to generate a database topology; associate the database topology to generate an association relationship for storing features; combine the database topology and association relationship to generate a detection profile; perform root cause analysis on the detection profile to obtain analysis results.

[0149] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0150] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0151] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0152] Receive cloud platform alarm events and parse them to obtain analysis results; load the database architecture from the cloud platform based on the analysis results to generate a database topology; associate the database topology to generate an association relationship for storing features; combine the database topology and association relationship to generate a detection profile; perform root cause analysis on the detection profile to obtain analysis results.

[0153] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0154] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0155] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0156] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0157] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0158] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A root cause analysis method for a database, characterized in that: include: Receive cloud platform alarm events and analyze them to obtain analysis results; Load the database schema from the cloud platform based on the parsing results to generate the database topology; Associating the database topology to generate association relationships for storing features; Combine database topology and relationships to generate detection profiles; Perform root cause analysis on the detection profile to obtain analysis results.

2. The root cause analysis method of a database according to claim 1, characterized in that: The receiving of cloud platform alarm events and parsing the alarm events to obtain parsing results include: Establish an alarm event receiving interface to receive alarm events from the cloud platform in real time or periodically; Analyze the content of received alarm events to extract alarm resources and alarm indicators; Based on the alarm indicator, search the resource classification database or configuration file for the corresponding resource category to determine the specific type of alarm resource; Query the detailed information of the alarm resource from the configuration management database; Integrate alarm resources, alarm indicators, resource categories and detailed information to obtain analytical results.

3. The root cause analysis method of a database according to claim 1, characterized in that: The database architecture is loaded from the cloud platform according to the parsing results to generate a database topology structure, including: Load the database architecture and operation status information through the API or management interface provided by the cloud platform; Use the link tracking system or the application management function of the cloud platform to identify and load the application information that has called the database in the recent period to obtain related applications; Load the architecture information and running status of the associated applications from the cloud platform according to the associated applications; Integrate the database's architecture information and operating status information, as well as the architecture information and operating status of associated applications, to build a database topology.

4. The root cause analysis method of a database according to claim 1, characterized in that: The associating the database topology structure to generate an association relationship for storing features includes: According to the analysis requirements, determine the features that need to be associated to obtain the feature set; Extracting detailed information of each feature from the database topology structure according to the feature set to obtain feature information data; Based on the feature information data, associations between features are established to generate association relationships for storing features.

5. The root cause analysis method of a database according to claim 1, characterized in that: The combination of database topology and association relationships to generate a detection profile includes: Collect the database's latest change records, maintenance operations, alarm events, service status, performance indicators, and number of connections from the database topology to obtain self-collected information; Collect relevant change records, maintenance operations, alarm events, status, and performance information from resources associated with the database based on the association relationships to obtain associated collection information; Use application performance management tools or logs to collect change records, release information, and database call performance information of upstream applications associated with the database to obtain application collection information. Integrate self-collected information, related collection information and application collection information to generate a detection profile.

6. The root cause analysis method of a database according to claim 1, characterized in that: The root cause analysis of the detection profile to obtain analysis results includes: Test the database itself, database-related resources, and upstream applications based on the test profile to obtain test results for the database itself, database-related resources, and upstream applications. Identify the detection portrait to get the apparent problem; Perform correlation analysis between the appearance problem and the database's own detection results to obtain the database correlation degree, and add a weighted edge between the appearance problem and the database's own detection results based on the database correlation degree, namely the database weight; Perform correlation analysis on the appearance problem and the database-related resource detection result to obtain the database-related resource correlation degree, and add a weighted edge between the appearance problem and the database-related resource detection result according to the database-related resource correlation degree, namely, the database-related resource weight; Perform correlation analysis between the appearance problem and the upstream application detection results to obtain the upstream application correlation degree. Based on the upstream application correlation degree, add a weighted edge between the appearance problem and the upstream application detection results, namely the upstream application weight. Integrate the surface issues, database test results, database-related resource test results, upstream application test results, database weights, database-related resource weights, and upstream application weights to construct a weight graph; The weight graph is analyzed to obtain analysis results.

7. The root cause analysis method of a database according to claim 1, characterized in that: After performing root cause analysis on the detection profile to obtain analysis results, the method further includes: Visualize the analysis results.

8. A root cause analysis device for a database, characterized in that: include: The receiving and parsing unit is used to receive the cloud platform alarm events and parse the alarm events to obtain the parsing results; A loading unit, used to load the database architecture from the cloud platform according to the parsing results to generate a database topology structure; An association unit, used to associate the database topology structure to generate an association relationship for storing features; A combining unit, used to combine the database topology and association relationships to generate a detection profile; The analysis unit is used to perform root cause analysis on the detection profile to obtain analysis results.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.